DeepSeek Guide — whale logoDeepSeek GuideFAN SITE
INTEGRATION GUIDE4 MIN READ

DeepSeek V4 Pro Responses API & Codex Setup

UPDATED: AUG 16, 2026AUTHOR: INDEPENDENT FAN GUIDE
OVERVIEW

DeepSeek V4 Pro natively supports OpenAI's Responses API: one-click Codex setup, models.json metadata, SSE stream — full guide.

01

What the Responses API Means

The OpenAI Responses API is the successor to Chat Completions — a stateful, tool-friendly surface used by Codex and modern agent clients. With the August 13 GA, DeepSeek V4 Pro natively supports it at the same base_url, which unlocks Codex as a first-class client for DeepSeek[1][2].

Before GA, Responses API on the official API worked only for Flash; V4 Pro requests errored out. The 0813 release flipped that switch, and DeepSeek now ships an official Codex integration guide[2].

NOTE

Timeline per DeepSeek GA notes and community reports around August 6-13, 2026[2].

02

Call It From the OpenAI SDK

Use the standard openai SDK with DeepSeek's base_url and the responses.create method. Model ID stays deepseek-v4-pro[1][2].

The missing [DONE] terminator trips up ports from Chat Completions: parsers that wait for it will hang. Listen for the response.completed event instead, then finalize from response.output[2].

Tool use inside Responses follows the same pattern as Chat Completions: define functions in tools, and the model returns function_call items in the output array. For a coding agent, pass instructions plus tools in one call and iterate on the returned items — the stateful conversation object keeps the thread across turns[2].

  • base_url stays https://api.deepseek.com — Responses API lives on the same root[1].
  • reasoning_effort maps to low/high/max as in Chat Completions[3].
  • Streaming: SSE events from response.created to response.completed — no data: [DONE] terminator[2].
NOTE

Streaming behavior per DeepSeek's Responses API guide[2].

example_code.py
from openai import OpenAI

client = OpenAI(
    api_key="<DeepSeek API Key>",
    base_url="https://api.deepseek.com",
)

resp = client.responses.create(
    model="deepseek-v4-pro",
    instructions="You are a coding agent. Return concise diffs.",
    input="Fix the failing test in tests/api_test.py",
    reasoning_effort="high",   # low / high / max
)
print(resp.output_text)
03

One-Click Codex Setup

DeepSeek ships a one-click script that backs up your ~/.codex config, writes a models.json with the correct V4 Pro/Flash metadata, and adds a deepseek provider block to config.toml. All Codex clients — CLI, ChatGPT desktop, and the VS Code extension — share that config, so you configure once[2].

  • Backs up existing config before touching anything[2].
  • Writes reasoning metadata: context_window 1048576, reasoning levels low/high/max, tool format[2].
  • Re-runnable: switch between Flash and Pro or restore your original config[2].
NOTE

Script behavior per DeepSeek's official Codex integration guide[2].

example_code.py
# macOS / Linux
bash <(curl -fsSL https://cdn.deepseek.com/api-docs/codex-deepseek-setup-en.sh)

# Windows PowerShell
irm https://cdn.deepseek.com/api-docs/codex-deepseek-setup-en.ps1 | iex

# The script backs up ~/.codex/config.toml, writes ~/.codex/models.json,
# adds [model_providers.deepseek], and validates. Re-run to switch models.
04

Manual models.json Configuration

If you prefer hand-rolled config, write the provider and model metadata yourself. The key fields Codex needs are the context window, reasoning levels, and tool-call format[2].

Then add the provider in ~/.codex/config.toml with model_providers.deepseek pointing at https://api.deepseek.com and set model_provider = "deepseek" in the model section[2].

NOTE

Field names follow Codex's models.json schema as documented by DeepSeek[2].

example_code.py
// ~/.codex/models.json (excerpt)
{
  "deepseek-v4-pro": {
    "context_window": 1048576,
    "reasoning_levels": ["low", "high", "max"],
    "default_reasoning_level": "high",
    "tool_format": "openai",
    "supports_streaming": true
  },
  "deepseek-v4-flash": {
    "context_window": 1048576,
    "reasoning_levels": ["low", "high", "max"],
    "default_reasoning_level": "high",
    "tool_format": "openai"
  }
}
05

Responses vs Chat Completions

Chat Completions remains fully supported and is the right choice for stateless, high-throughput pipelines. Responses API adds a stateful conversation model with native tool orchestration — the surface Codex and next-gen agent clients expect[1][2].

Both routes bill at the same per-token rates and share prefix caching[4]. If you are building a coding agent today, start on Responses API; if you are migrating a chat pipeline, Chat Completions keeps working untouched[1].

A practical migration path: keep your existing Chat Completions code for batch jobs, and add a Responses API branch only for the agent path (Codex, IDE assistants). The two surfaces share the same model and key, so you can run them side by side during the transition and retire the older surface when your agent code is stable[1][2].

FactorChat CompletionsResponses API
StatefulnessStatelessStateful (conversation object)
Tool orchestrationManualNative (Codex-style)
Codex supportVia compatibilityFirst-class[2]
Best forBatch, ingestion, chatAgent loops, Codex, IDE agents
NOTE

Feature split per DeepSeek's Responses API and quick-start docs[1][2].

Sponsored
Sponsored