> ## Documentation Index
> Fetch the complete documentation index at: https://docs.socialvision.tisyk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# ChatGPT 兼容对话 (Chat)

> 统一 OpenAI 兼容接口，支持 GPT-5.6、GPT-6 Astra、o1/o3-mini、Claude 5.5、Gemini 3.8 等全系模型

## 对话补全（OpenAI 兼容）

`POST /v1/chat/completions`

统一 OpenAI 兼容入口。通过 `model` 选择上游，例如：

* `gpt-5.6` / `gpt-6-astra` / `gpt-4o` → OpenAI
* `claude-5.5-sonnet` / `claude-fable-5.1` → Anthropic
* `gemini-3.8-pro` / `gemini-3.8-flash` → Google Gemini
* 其他模型见 `GET /v1/models`

OpenAI 官方文档：[https://platform.openai.com/docs/api-reference/chat/create](https://platform.openai.com/docs/api-reference/chat/create)

模型名称格式为：gpt-4-gizmo-\*，系统会自动进行识别

比如这个GPTs：[https://chatgpt.com/g/g-B3hgivKK9-write-for-me](https://chatgpt.com/g/g-B3hgivKK9-write-for-me)

那么它的模型名应该填写为：gpt-4-gizmo-g-B3hgivKK9

GPTs列表：[https://chatgpt.com/gpts](https://chatgpt.com/gpts)

### 请求体参数 (Body)

<ParamField body="enable_thinking" type="boolean">
  是否开启思考模式（本仓库扩展）。仅布尔值 true 生效；与 extra\_body.enable\_thinking 二选一即可。
</ParamField>

<ParamField body="extra_body" type="object">
  对话扩展参数。常用：enable\_thinking；部分 Gemini 模型可用 google.thinking\_config 等。是否生效取决于渠道与模型。
</ParamField>

<ParamField body="frequency_penalty" type="number">
  频率惩罚，约 -2～2。正值降低重复相同词句的概率。
</ParamField>

<ParamField body="logit_bias" type="object">
  按 token ID 调整出现概率，取值为 -100～100 的映射对象。
</ParamField>

<ParamField body="max_completion_tokens" type="integer">
  补全 token 上限（o 系、GPT-5 等推理模型更常用）。未设置时可能沿用 max\_tokens。
</ParamField>

<ParamField body="max_tokens" type="integer">
  本次回复最多生成的 token 数（常用字段）。
</ParamField>

<ParamField body="messages" type="array" required>
  对话消息列表，至少一条。支持多轮：system / user / assistant / tool。
</ParamField>

<ParamField body="model" type="string" required>
  要使用的模型 ID。网关会按令牌分组选择渠道，并可能做模型名映射或后缀解析（如 -thinking、-search）。
</ParamField>

<ParamField body="n" type="integer">
  对同一轮输入生成几条 assistant 回复候选。
</ParamField>

<ParamField body="parallel_tool_calls" type="boolean">
  是否允许在一次回复中并行发起多个工具调用（上游支持时）。
</ParamField>

<ParamField body="presence_penalty" type="number">
  存在惩罚，约 -2～2。正值减少重复已出现过的主题表述。
</ParamField>

<ParamField body="reasoning_effort" type="string">
  推理模型力度（如 low / medium / high）。也可通过模型名后缀由网关自动注入。
</ParamField>

<ParamField body="response_format" type="object" />

<ParamField body="seed" type="number">
  随机种子，在相同参数下尽量得到可复现结果（不保证完全一致）。注意字段名为 seed。
</ParamField>

<ParamField body="stop" type="any">
  遇到这些字符串时停止生成。可为单个字符串或字符串数组（最多 4 个）。
</ParamField>

<ParamField body="stream" type="boolean">
  是否流式输出。false：一次返回完整 JSON；true：SSE 增量推送，结束为 data: \[DONE]。
</ParamField>

<ParamField body="stream_options" type="object" />

<ParamField body="temperature" type="number">
  采样温度，通常 0～2。值越高回复越随机，越低越稳定。与 top\_p 一般只调一个。
</ParamField>

<ParamField body="tool_choice" type="any">
  控制是否调用工具：none（不调用）、auto（自动）、required（必须调用），或指定某个函数。
</ParamField>

<ParamField body="tools" type="array">
  模型可调用的工具（函数）列表，用于对话中的 function calling。
</ParamField>

<ParamField body="top_p" type="number">
  核采样，0～1。只从累计概率达到 top\_p 的 token 中采样。
</ParamField>

<ParamField body="user" type="string">
  终端用户唯一标识，便于平台侧滥用检测与统计。
</ParamField>

<ParamField body="web_search_options" type="object" />

### 请求示例 (JSON)

```json theme={null}
{
  "extra_body": {
    "enable_thinking": true
  },
  "messages": [
    {
      "content": "有三个盒子，只有一个里面有奖品。你选 2 号，主持人打开 3 号（空），问要不要改选 1 号？分析并给建议。",
      "role": "user"
    }
  ],
  "model": "deepseek-chat",
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}
```

### 请求示例 (cURL)

```bash theme={null}
curl -X POST "https://socialvision.tisyk.xyz/v1/chat/completions" \
  -H "Authorization: Bearer sk-YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "extra_body": {
    "enable_thinking": true
  },
  "messages": [
    {
      "content": "有三个盒子，只有一个里面有奖品。你选 2 号，主持人打开 3 号（空），问要不要改选 1 号？分析并给建议。",
      "role": "user"
    }
  ],
  "model": "deepseek-chat",
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}'
```

### 响应示例 (200 OK)

```json theme={null}
{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "logprobs": "",
      "message": {
        "annotations": [],
        "content": "\n\nHello there, how may I assist you today?",
        "refusal": "",
        "role": "assistant"
      }
    }
  ],
  "created": 1677652288,
  "id": "chatcmpl-123",
  "model": "gpt-4o",
  "object": "chat.completion",
  "service_tier": "default",
  "system_fingerprint": "fp_abc123",
  "usage": {
    "completion_tokens": 12,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "latency_checkpoint": {
      "engine_tbt_ms": 0,
      "engine_ttft_ms": 0,
      "engine_ttlt_ms": 0,
      "pre_inference_ms": 0,
      "service_tbt_ms": 0,
      "service_ttft_ms": 0,
      "service_ttlt_ms": 0,
      "total_duration_ms": 0,
      "user_visible_ttft_ms": 0
    },
    "prompt_tokens": 9,
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    },
    "total_tokens": 21
  }
}
```

***

## 对话补全（OpenAI 兼容）

`POST /v1/chat/completions`

统一 OpenAI 兼容入口。通过 `model` 选择上游，例如：

* `gpt-4o` / `gpt-4o-audio-preview` → OpenAI
* `gemini-2.5-flash-all` → Google Gemini（兼容格式）
* 其他模型见 `GET /v1/models`

OpenAI 官方文档：[https://platform.openai.com/docs/api-reference/chat/create](https://platform.openai.com/docs/api-reference/chat/create)

模型名称格式为：gpt-4-gizmo-\*，系统会自动进行识别

比如这个GPTs：[https://chatgpt.com/g/g-B3hgivKK9-write-for-me](https://chatgpt.com/g/g-B3hgivKK9-write-for-me)

那么它的模型名应该填写为：gpt-4-gizmo-g-B3hgivKK9

GPTs列表：[https://chatgpt.com/gpts](https://chatgpt.com/gpts)

### 请求体参数 (Body)

<ParamField body="enable_thinking" type="boolean">
  是否开启思考模式（本仓库扩展）。仅布尔值 true 生效；与 extra\_body.enable\_thinking 二选一即可。
</ParamField>

<ParamField body="extra_body" type="object">
  对话扩展参数。常用：enable\_thinking；部分 Gemini 模型可用 google.thinking\_config 等。是否生效取决于渠道与模型。
</ParamField>

<ParamField body="frequency_penalty" type="number">
  频率惩罚，约 -2～2。正值降低重复相同词句的概率。
</ParamField>

<ParamField body="logit_bias" type="object">
  按 token ID 调整出现概率，取值为 -100～100 的映射对象。
</ParamField>

<ParamField body="max_completion_tokens" type="integer">
  补全 token 上限（o 系、GPT-5 等推理模型更常用）。未设置时可能沿用 max\_tokens。
</ParamField>

<ParamField body="max_tokens" type="integer">
  本次回复最多生成的 token 数（常用字段）。
</ParamField>

<ParamField body="messages" type="array" required>
  对话消息列表，至少一条。支持多轮：system / user / assistant / tool。
</ParamField>

<ParamField body="model" type="string" required>
  要使用的模型 ID。网关会按令牌分组选择渠道，并可能做模型名映射或后缀解析（如 -thinking、-search）。
</ParamField>

<ParamField body="n" type="integer">
  对同一轮输入生成几条 assistant 回复候选。
</ParamField>

<ParamField body="parallel_tool_calls" type="boolean">
  是否允许在一次回复中并行发起多个工具调用（上游支持时）。
</ParamField>

<ParamField body="presence_penalty" type="number">
  存在惩罚，约 -2～2。正值减少重复已出现过的主题表述。
</ParamField>

<ParamField body="reasoning_effort" type="string">
  推理模型力度（如 low / medium / high）。也可通过模型名后缀由网关自动注入。
</ParamField>

<ParamField body="response_format" type="object" />

<ParamField body="seed" type="number">
  随机种子，在相同参数下尽量得到可复现结果（不保证完全一致）。注意字段名为 seed。
</ParamField>

<ParamField body="stop" type="any">
  遇到这些字符串时停止生成。可为单个字符串或字符串数组（最多 4 个）。
</ParamField>

<ParamField body="stream" type="boolean">
  是否流式输出。false：一次返回完整 JSON；true：SSE 增量推送，结束为 data: \[DONE]。
</ParamField>

<ParamField body="stream_options" type="object" />

<ParamField body="temperature" type="number">
  采样温度，通常 0～2。值越高回复越随机，越低越稳定。与 top\_p 一般只调一个。
</ParamField>

<ParamField body="tool_choice" type="any">
  控制是否调用工具：none（不调用）、auto（自动）、required（必须调用），或指定某个函数。
</ParamField>

<ParamField body="tools" type="array">
  模型可调用的工具（函数）列表，用于对话中的 function calling。
</ParamField>

<ParamField body="top_p" type="number">
  核采样，0～1。只从累计概率达到 top\_p 的 token 中采样。
</ParamField>

<ParamField body="user" type="string">
  终端用户唯一标识，便于平台侧滥用检测与统计。
</ParamField>

<ParamField body="web_search_options" type="object" />

### 请求示例 (JSON)

```json theme={null}
{
  "extra_body": {
    "enable_thinking": true
  },
  "messages": [
    {
      "content": "有三个盒子，只有一个里面有奖品。你选 2 号，主持人打开 3 号（空），问要不要改选 1 号？分析并给建议。",
      "role": "user"
    }
  ],
  "model": "deepseek-chat",
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}
```

### 请求示例 (cURL)

```bash theme={null}
curl -X POST "https://socialvision.tisyk.xyz/v1/chat/completions" \
  -H "Authorization: Bearer sk-YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "extra_body": {
    "enable_thinking": true
  },
  "messages": [
    {
      "content": "有三个盒子，只有一个里面有奖品。你选 2 号，主持人打开 3 号（空），问要不要改选 1 号？分析并给建议。",
      "role": "user"
    }
  ],
  "model": "deepseek-chat",
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}'
```

### 响应示例 (200 OK)

```json theme={null}
{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "logprobs": "",
      "message": {
        "annotations": [],
        "content": "\n\nHello there, how may I assist you today?",
        "refusal": "",
        "role": "assistant"
      }
    }
  ],
  "created": 1677652288,
  "id": "chatcmpl-123",
  "model": "gpt-4o",
  "object": "chat.completion",
  "service_tier": "default",
  "system_fingerprint": "fp_abc123",
  "usage": {
    "completion_tokens": 12,
    "completion_tokens_details": {
      "accepted_prediction_tokens": 0,
      "audio_tokens": 0,
      "reasoning_tokens": 0,
      "rejected_prediction_tokens": 0
    },
    "latency_checkpoint": {
      "engine_tbt_ms": 0,
      "engine_ttft_ms": 0,
      "engine_ttlt_ms": 0,
      "pre_inference_ms": 0,
      "service_tbt_ms": 0,
      "service_ttft_ms": 0,
      "service_ttlt_ms": 0,
      "total_duration_ms": 0,
      "user_visible_ttft_ms": 0
    },
    "prompt_tokens": 9,
    "prompt_tokens_details": {
      "audio_tokens": 0,
      "cached_tokens": 0
    },
    "total_tokens": 21
  }
}
```

***

## 创建完成

`POST /v1/completions`

给定一个提示，该模型将返回一个或多个预测的完成，并且还可以返回每个位置的替代标记的概率。

为提供的提示和参数创建完成

[https://platform.openai.com/docs/api-reference/completions](https://platform.openai.com/docs/api-reference/completions)

### 请求参数 (Query / Path)

<ParamField header="Authorization" type="string">
  Bearer API Key
</ParamField>

### 请求体参数 (Body)

<ParamField body="best_of" type="integer">
  默认为1 在服务器端生成best\_of个补全,并返回“最佳”补全(每个令牌的日志概率最高的那个)。无法流式传输结果。  与n一起使用时,best\_of控制候选补全的数量,n指定要返回的数量 – best\_of必须大于n。  注意:因为这个参数会生成许多补全,所以它可以快速消耗您的令牌配额。请谨慎使用,并确保您对max\_tokens和stop有合理的设置。
</ParamField>

<ParamField body="echo" type="boolean">
  默认为false 除了补全之外,还回显提示
</ParamField>

<ParamField body="frequency_penalty" type="number">
  默认为0 -2.0和2.0之间的数字。正值根据文本目前的现有频率处罚新令牌,降低模型逐字重复相同行的可能性。
</ParamField>

<ParamField body="logit_bias" type="object">
  默认为null 修改完成中指定令牌出现的可能性。  接受一个JSON对象,该对象将令牌(由GPT令牌化器中的令牌ID指定)映射到关联偏差值,-100到100。您可以使用这个令牌化器工具(适用于GPT-2和GPT-3)将文本转换为令牌ID。从数学上讲,偏差在对模型进行采样之前添加到生成的logit中。确切效果因模型而异,但-1至1之间的值应降低或提高选择的可能性;像-100或100这样的值应导致相关令牌的禁用或专属选择。  例如,您可以传递\{"50256": -100}来防止生成\<|endoftext|>令牌。
</ParamField>

<ParamField body="logprobs" type="string">
  默认为null
  包括logprobs个最可能令牌的日志概率,以及所选令牌。例如,如果logprobs为5,API将返回5个最有可能令牌的列表。 API总会返回采样令牌的logprob,因此响应中最多可能有logprobs+1个元素。

  logprobs的最大值是5。
</ParamField>

<ParamField body="max_tokens" type="integer">
  默认为16
  在补全中生成的最大令牌数。

  提示的令牌计数加上max\_tokens不能超过模型的上下文长度。 计数令牌的Python代码示例。
</ParamField>

<ParamField body="model" type="string" required>
  要使用的模型的 ID。您可以使用[List models](https://platform.openai.com/docs/api-reference/models/list) API 来查看所有可用模型，或查看我们的[模型概述](https://platform.openai.com/docs/models/overview)以了解它们的描述。
</ParamField>

<ParamField body="n" type="integer">
  默认为1
  为每个提示生成的补全数量。

  注意:因为这个参数会生成许多补全,所以它可以快速消耗您的令牌配额。请谨慎使用,并确保您对max\_tokens和stop有合理的设置。
</ParamField>

<ParamField body="presence_penalty" type="number">
  默认为0 -2.0和2.0之间的数字。正值根据它们是否出现在目前的文本中来惩罚新令牌,增加模型讨论新话题的可能性。  有关频率和存在惩罚的更多信息,请参阅。
</ParamField>

<ParamField body="prompt" type="string" required>
  生成完成的提示，编码为字符串、字符串数组、标记数组或标记数组数组。  请注意，\<|endoftext|> 是模型在训练期间看到的文档分隔符，因此如果未指定提示，模型将生成新文档的开头。
</ParamField>

<ParamField body="seed" type="integer">
  如果指定,我们的系统将尽最大努力确定性地进行采样,以便使用相同的种子和参数的重复请求应返回相同的结果。  不保证确定性,您应该参考system\_fingerprint响应参数来监视后端的更改。
</ParamField>

<ParamField body="stop" type="string">
  默认为null 最多4个序列,API将停止在其中生成更多令牌。返回的文本不会包含停止序列。
</ParamField>

<ParamField body="stream" type="boolean">
  默认为false 是否流回部分进度。如果设置,令牌将作为可用时发送为仅数据的服务器发送事件,流由数据 Terminated by a data: \[DONE] message. 对象消息终止。 Python代码示例。
</ParamField>

<ParamField body="suffix" type="string">
  默认为null 在插入文本的补全之后出现的后缀。
</ParamField>

<ParamField body="temperature" type="integer">
  默认为1 要使用的采样温度,介于0和2之间。更高的值(如0.8)将使输出更随机,而更低的值(如0.2)将使其更集中和确定。  我们通常建议更改这个或top\_p,而不是两者都更改。
</ParamField>

<ParamField body="top_p" type="integer">
  表示最终用户的唯一标识符,这可以帮助OpenAI监控和检测滥用。 了解更多。
</ParamField>

<ParamField body="user" type="string" required>
  终端用户标识（风控/计费辅助）
</ParamField>

### 请求示例 (JSON)

```json theme={null}
{
  "max_tokens": 30,
  "model": "gpt-3.5-turbo-instruct",
  "prompt": "你好,",
  "temperature": 0,
  "user": "default-user"
}
```

### 请求示例 (cURL)

```bash theme={null}
curl -X POST "https://socialvision.tisyk.xyz/v1/completions" \
  -H "Authorization: Bearer sk-YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "max_tokens": 30,
  "model": "gpt-3.5-turbo-instruct",
  "prompt": "你好,",
  "temperature": 0,
  "user": "default-user"
}'
```

### 响应示例 (200 OK)

```json theme={null}
{
  "choices": [
    {
      "finish_reason": "length",
      "index": 0,
      "logprobs": "",
      "text": "我是小冰,很高兴认识你。我是一个人工智能助手，可以回"
    }
  ],
  "created": 1753859563,
  "id": "cmpl-ByvHP6AWeB1L5vWZSPNHsB12sU9db",
  "model": "gpt-3.5-turbo-instruct",
  "object": "text_completion",
  "system_fingerprint": "fp_abc123",
  "usage": {
    "completion_tokens": 30,
    "prompt_tokens": 3,
    "total_tokens": 33
  }
}
```

***


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.