コンテンツにスキップ

チャット補完

POST /v1/chat/completions

指定した会話に対してモデル応答を生成します。OpenAI Chat Completions API と完全互換です。

パラメータ型必須説明
modelstringはいモデル ID (モデル を参照)
messagesarrayはい会話メッセージ (メッセージタイプ を参照)
streambooleanいいえSSE で部分デルタをストリーミング。デフォルト: false
stream_optionsobjectいいえ{ "include_usage": true } で最終ストリームイベントにトークン数を含める
temperaturenumberいいえサンプリング温度 (0〜2)。デフォルトはモデル依存
top_pnumberいいえNucleus サンプリングの閾値 (0〜1)
nintegerいいえ生成する選択肢の数 (1〜128)
max_tokensintegerいいえ生成する最大トークン数 (非推奨 — max_completion_tokens を使用)
max_completion_tokensintegerいいえ推論トークンを含む生成トークン数の上限
stopstring | string[]いいえ最大 4 つの停止シーケンス
frequency_penaltynumberいいえ頻度ペナルティ (-2〜2)
presence_penaltynumberいいえ出現ペナルティ (-2〜2)
logprobsbooleanいいえ出力トークンの対数確率を返す
top_logprobsintegerいいえ位置ごとに最も可能性の高いトークン数 (0〜20)
logit_biasobjectいいえトークン ID からバイアス値 (-100〜100) へのマップ
response_formatobjectいいえ{ "type": "text" }、{ "type": "json_object" }、または { "type": "json_schema", "json_schema": {...} }
seedintegerいいえ決定論的サンプリングのためのシード
toolsarrayいいえモデルが呼び出し可能な関数ツール
tool_choicestring | objectいいえ"none"、"auto"、"required"、または特定のツール
parallel_tool_callsbooleanいいえ並列関数呼び出しを許可
reasoning_effortstringいいえモデルごとに異なります。GET /v1/models の reasoning_efforts を参照してください。モデルが受け付けない値は拒否されます
top_kintegerいいえTop-k サンプリング (プロバイダ依存)
min_pnumberいいえMin-p サンプリングの閾値 (0〜1、プロバイダ依存)
repetition_penaltynumberいいえ繰り返しペナルティ (プロバイダ依存)
userstringいいえ不正利用追跡用のエンドユーザー識別子
{ "role": "system", "content": "You are a helpful assistant." }
{ "role": "user", "content": "What is the capital of France?" }

ユーザーメッセージはマルチモーダルなコンテンツ配列も受け付けます:

{
"role": "user",
"content": [
{ "type": "text", "text": "What's in this image?" },
{ "type": "image_url", "image_url": { "url": "https://...", "detail": "auto" } }
]
}

対応するコンテンツタイプ: text、image_url、video_url、audio_url、input_audio、file。

画像については、インラインの base64 ではなく Files API でアップロードして file_id で参照することを推奨します。マルチターンの会話やリトライ時に、クライアントからバイト列を再送する必要がなくなります:

{
"role": "user",
"content": [
{ "type": "text", "text": "What's in this image?" },
{ "type": "file", "file": { "file_id": "file-abc123" } }
]
}

モデルは対応する capability (例: vision) を持つ必要があります。持たない場合、リクエストは 400 model_capability_mismatch で拒否されます。

{ "role": "assistant", "content": "The capital of France is Paris." }
{ "role": "tool", "tool_call_id": "call_abc123", "content": "{\"result\": 42}" }
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1700000000,
"model": "deepseek-ai/deepseek-v4-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 9,
"total_tokens": 19
}
}
フィールド型説明
prompt_tokensinteger入力で消費したトークン数
completion_tokensinteger生成された出力トークン数
total_tokensinteger入力と出力の合計
prompt_tokens_detailsobject任意。{ cached_tokens, audio_tokens }
completion_tokens_detailsobject任意。{ reasoning_tokens, audio_tokens }

stream: true を指定すると、Server-Sent Events として部分的な応答を受け取れます。

Terminal window
curl https://api.aiand.com/v1/chat/completions \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-ai/deepseek-v4-flash",
"stream": true,
"messages": [{"role": "user", "content": "Count to 5"}]
}'

各イベントは JSON チャンクを含む data: 行で構成されます。ストリームは data: [DONE] で終了します。

最終イベントにトークン使用量を含めるには:

{
"stream": true,
"stream_options": { "include_usage": true }
}

リクエストでツールを定義すると、モデルがそれらを呼び出すかを選択できます:

{
"model": "deepseek-ai/deepseek-v4-flash",
"messages": [{ "role": "user", "content": "What's the weather in Tokyo?" }],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": { "type": "string" }
},
"required": ["location"]
}
}
}
]
}

モデルがツールを呼び出すと、レスポンスには tool_calls が含まれます:

{
"choices": [
{
"message": {
"role": "assistant",
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\": \"Tokyo\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
]
}

実行結果は tool メッセージとして返送し、会話を継続します。