आमंत्रित करें और कमाएँ

आमंत्रण पुरस्कार कैसे काम करते हैं

अपना आमंत्रण लिंक साझा करें। मित्र इसके माध्यम से पंजीकरण करके टॉप-अप करता है तो उसके बाद के टॉप-अप पर आपको दिखाया गया पुरस्कार मिलेगा।

Anthropic Messages, Chat Completions और Responses: चयन, रूपांतरण और compatibility की सीमाएँ

HTTP 200 मिलने का अर्थ यह नहीं कि Agent compatible है। तीन पूरे tool-call round trips के जरिए Messages, Chat Completions और Responses के वास्तविक अंतर, migration, testing और troubleshooting समझें।

विषय-सूची

एक ही Agent में endpoint और field names बदलने के बाद request 200 लौटा सकती है, फिर भी टूल execute न हो, JSON Schema से मेल न खाए या multi-turn context अचानक खो जाए। इसका अर्थ आम तौर पर यह नहीं कि मॉडल “कम बुद्धिमान” हो गया है। समस्या यह होती है कि application ने तीन अलग protocols को एक ही interface मान लिया।

किसी migration को सफल मानने के लिए कम से कम तीन स्तर जाँचने चाहिए:

  1. Format स्वीकार होता है: server request को parse कर सकता है और success status लौटाता है।
  2. Behavior समान रहता है: टूल call होते हैं, results सही तरीके से वापस जाते हैं, stream पूरी तरह समाप्त होती है और multi-turn context जुड़ा रहता है।
  3. Capabilities सुरक्षित रहती हैं: strict Schema, native reasoning state, hosted tools, structured output जैसी क्षमताएँ silently ignore या downgrade नहीं होतीं।

HTTP 200 केवल पहले स्तर को सिद्ध करता है। केवल एक text block बनाने वाली request के लिए साधारण conversion पर्याप्त हो सकता है। लेकिन Agent, tool calls, streamed arguments, multi-turn state या reasoning models के साथ पूरी interaction chain को validate करना आवश्यक है।

व्यावहारिक निष्कर्ष: protocol को client और आवश्यक capabilities के अनुसार चुनें, model name के अनुसार नहीं

परिस्थितिबेहतर शुरुआती विकल्पकारण
Existing application OpenAI SDK और messages को स्थिर रूप से उपयोग करती हैChat Completionsसबसे कम बदलाव; मौजूदा message और tool loop जारी रह सकती है
नया OpenAI Agent hosted tools, typed Items या server-side state continuation चाहता हैResponsesOpenAI फिलहाल नए projects के लिए इसकी सिफारिश करता है और Agent capabilities अधिक व्यापक हैं
Claude Code, native Claude application या Claude-specific capabilities पर निर्भर workflowAnthropic MessagesContent blocks, tool results, thinking और संबंधित behavior Anthropic के native contract का पालन करते हैं
अपना gateway या multi-model routerहर upstream protocol के लिए अलग adapter रखेंएक “universal JSON” सभी native capabilities को बिना हानि व्यक्त नहीं कर सकता

OpenAI अभी भी Chat Completions को support करता है, इसलिए स्थिर production application को केवल नई interface उपलब्ध होने के कारण तुरंत दोबारा लिखने की जरूरत नहीं है। नई परियोजना या Responses-native capabilities की आवश्यकता होने पर migration अधिक उचित है। Anthropic Messages भी केवल OpenAI interface नहीं है जिसमें field का नाम बदलकर messages कर दिया गया हो। इसके content blocks, tool handoff, stream events और state rules अलग contract बनाते हैं।

तीनों API के मुख्य अंतर

इस लेख में Completions से आशय Chat Completions है, पुराने /v1/completions endpoint से नहीं।

आयामOpenAI Chat CompletionsOpenAI ResponsesAnthropic Messages
Endpoint/v1/chat/completions/v1/responses/v1/messages
मुख्य inputmessagesinput Items; सरल message input भी स्वीकारmessages, सामान्यतः अलग top-level system के साथ
मुख्य outputchoices[].messageoutput[] में typed Itemscontent[] में content blocks
Tool definitiontools[].functiontools[] में सीधे name और parameterstools[] में input_schema
Tool argumentsfunction.arguments, JSON stringarguments, JSON stringtool_use.input, JSON object
Correlation IDtool_calls[].idcall_idtool_use.id
Tool result वापस देनाrole: "tool" + tool_call_idfunction_call_output + call_iduser message में tool_result + tool_use_id
Multi-turn stateApplication message history को replay करती हैItems replay, previous_response_id या ConversationsApplication messages और content blocks replay करती है
Final structured outputresponse_formattext.formatoutput_config.format
Streamingchoices[].deltaTyped Responses eventsmessage/content block events

तालिका देखकर केवल field names बदलने का मामला लग सकता है। वास्तविक failure आम तौर पर दूसरी request में होता है: मॉडल के tool call लौटाने के बाद application उसे कैसे execute करे, कौन-सा ID बचाए और किस role तथा क्रम में result वापस भेजे? नीचे side effects से मुक्त एक ही task को तीनों protocols में पूरा चलाया गया है।

साझा उदाहरण: परीक्षण प्लान की जानकारी लेना

User का प्रश्न है:

team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।

Tool का नाम get_plan_info है। यह केवल निश्चित local data पढ़ता है और कोई external side effect नहीं बनाता, इसलिए protocol migration test के लिए उपयुक्त है।

नीचे दिया गया plan data शिक्षण के लिए synthetic है। यह OpenAI, Anthropic या BetterToken के वास्तविक plans, prices या entitlements का प्रतिनिधित्व नहीं करता। तीनों request-response sequences केवल protocol structure दिखाती हैं; ये live API execution records नहीं हैं।

Application-side tool को protocol-independent function के रूप में लिखा जा सकता है:

from __future__ import annotations

import json
from typing import Any


PLAN_FIXTURES: dict[str, dict[str, Any]] = {
    "team": {
        "plan_code": "team",
        "display_name": "Team",
        "billing_mode": "usage_based",
        "included_requests": 10_000,
        "overage_allowed": True,
        "source_version": "fixture-2026-09-01",
    }
}


def execute_tool(name: str, raw_arguments: str | dict[str, Any]) -> str:
    """शिक्षण के लिए बनाए गए read-only टूल को चलाता है और ऐसी JSON स्ट्रिंग लौटाता है जिसे सीधे मॉडल को दिया जा सके।"""
    if isinstance(raw_arguments, str):
        arguments = json.loads(raw_arguments)
    elif isinstance(raw_arguments, dict):
        arguments = raw_arguments
    else:
        raise TypeError("टूल arguments JSON स्ट्रिंग या object होने चाहिए")

    if name != "get_plan_info":
        raise ValueError(f"अज्ञात टूल: {name}")

    if set(arguments) != {"plan_code"}:
        raise ValueError("get_plan_info केवल plan_code स्वीकार करता है")

    plan_code = arguments["plan_code"]
    if not isinstance(plan_code, str):
        raise TypeError("plan_code एक स्ट्रिंग होना चाहिए")

    plan = PLAN_FIXTURES.get(plan_code)
    if plan is None:
        return json.dumps(
            {"ok": False, "error": "plan_not_found", "plan_code": plan_code},
            ensure_ascii=False,
        )

    return json.dumps({"ok": True, "data": plan}, ensure_ascii=False)

Request में strict Schema लगाने के बाद भी application को अपनी input validation रखनी चाहिए। Strict mode मॉडल से बने tool arguments को नियंत्रित करता है; यह authorization, enum validity, idempotency और business security checks का स्थान नहीं लेता।

Chat Completions: पूरा tool round trip

पहली request: मॉडल से tool call बनवाना

नीचे protocol समझाने के लिए OpenAI का official endpoint उपयोग किया गया है। Compatible service से जुड़ते समय provider documentation के अनुसार Base URL, authentication method और Model ID बदलें।

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_OPENAI_MODEL",
    "messages": [
      {
        "role": "system",
        "content": "आप प्लान संबंधी सहायक हैं। केवल टूल से लौटे डेटा के आधार पर उत्तर दें; अनुमान न लगाएँ।"
      },
      {
        "role": "user",
        "content": "team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।"
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_plan_info",
          "description": "प्लान कोड से निश्चित परीक्षण डेटा प्राप्त करें",
          "strict": true,
          "parameters": {
            "type": "object",
            "properties": {
              "plan_code": {
                "type": "string",
                "enum": ["team"]
              }
            },
            "required": ["plan_code"],
            "additionalProperties": false
          }
        }
      }
    ],
    "tool_choice": "required",
    "parallel_tool_calls": false
  }'

Application को assistant message के tool_calls पढ़ने हैं। नीचे केवल आगे के flow के लिए आवश्यक fields रखे गए हैं:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_plan_001",
            "type": "function",
            "function": {
              "name": "get_plan_info",
              "arguments": "{\"plan_code\":\"team\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}

दो values नहीं खोनी चाहिए:

  • tool_calls[0].id: दूसरी request में इसे बिना बदले tool_call_id के रूप में वापस भेजें।
  • function.arguments: यह JSON string है। पहले parse करें, फिर अपना Schema और business validation लगाएँ।

Tool execute करें:

tool_result = execute_tool(
    "get_plan_info",
    "{\"plan_code\":\"team\"}",
)

दूसरी request: tool result मॉडल को वापस देना

Chat Completions में पहले response से मिला assistant tool-call message history में रखना होता है, और उसके बाद role: "tool" वाला result message जोड़ना होता है।

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_OPENAI_MODEL",
    "messages": [
      {
        "role": "system",
        "content": "आप प्लान संबंधी सहायक हैं। केवल टूल से लौटे डेटा के आधार पर उत्तर दें; अनुमान न लगाएँ।"
      },
      {
        "role": "user",
        "content": "team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।"
      },
      {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_plan_001",
            "type": "function",
            "function": {
              "name": "get_plan_info",
              "arguments": "{\"plan_code\":\"team\"}"
            }
          }
        ]
      },
      {
        "role": "tool",
        "tool_call_id": "call_plan_001",
        "content": "{\"ok\":true,\"data\":{\"plan_code\":\"team\",\"display_name\":\"Team\",\"billing_mode\":\"usage_based\",\"included_requests\":10000,\"overage_allowed\":true,\"source_version\":\"fixture-2026-09-01\"}}"
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_plan_info",
          "description": "प्लान कोड से निश्चित परीक्षण डेटा प्राप्त करें",
          "strict": true,
          "parameters": {
            "type": "object",
            "properties": {
              "plan_code": {
                "type": "string",
                "enum": ["team"]
              }
            },
            "required": ["plan_code"],
            "additionalProperties": false
          }
        }
      }
    ]
  }'

एक representative final message:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "Team प्लान शामिल सीमा से अधिक उपयोग पर usage-based billing को समर्थन देता है। परीक्षण डेटा में 10,000 requests शामिल हैं और overage_allowed true है।"
      },
      "finish_reason": "stop"
    }
  ]
}

यदि adapter केवल पहली user request convert करता है, assistant के tool_calls नहीं बचाता, या गलत ID को tool_call_id में डालता है, तो दूसरी request उसी tool call की continuation नहीं रहती।

Responses: पूरा tool round trip

Responses messages, reasoning, tool calls और tool results को अलग-अलग Item types में दर्शाता है। output[0] को हर बार final text न मानें; प्रत्येक Item के type के अनुसार processing करें।

पहली request: मॉडल से function_call Item लेना

curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_OPENAI_MODEL",
    "instructions": "आप प्लान संबंधी सहायक हैं। केवल टूल से लौटे डेटा के आधार पर उत्तर दें; अनुमान न लगाएँ।",
    "input": "team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।",
    "tools": [
      {
        "type": "function",
        "name": "get_plan_info",
        "description": "प्लान कोड से निश्चित परीक्षण डेटा प्राप्त करें",
        "strict": true,
        "parameters": {
          "type": "object",
          "properties": {
            "plan_code": {
              "type": "string",
              "enum": ["team"]
            }
          },
          "required": ["plan_code"],
          "additionalProperties": false
        }
      }
    ],
    "tool_choice": "required",
    "parallel_tool_calls": false,
    "store": false
  }'

एक representative tool-call Item:

{
  "id": "resp_plan_001",
  "object": "response",
  "output": [
    {
      "type": "function_call",
      "id": "fc_plan_001",
      "call_id": "call_plan_001",
      "name": "get_plan_info",
      "arguments": "{\"plan_code\":\"team\"}",
      "status": "completed"
    }
  ]
}

Tool result को जोड़ने के लिए call_id उपयोग करें। id: "fc_plan_001" उस Item का अपना ID है; इसे call_id के स्थान पर उपयोग नहीं करना चाहिए।

Tool execute करें:

tool_result = execute_tool(
    "get_plan_info",
    "{\"plan_code\":\"team\"}",
)

दूसरी request: function_call_output वापस देना

नीचे stateless manual replay उपयोग किया गया है, इसलिए instructions, मूल user input, tool call और tool result फिर से भेजे जाते हैं।

curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_OPENAI_MODEL",
    "instructions": "आप प्लान संबंधी सहायक हैं। केवल टूल से लौटे डेटा के आधार पर उत्तर दें; अनुमान न लगाएँ।",
    "input": [
      {
        "role": "user",
        "content": "team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।"
      },
      {
        "type": "function_call",
        "call_id": "call_plan_001",
        "name": "get_plan_info",
        "arguments": "{\"plan_code\":\"team\"}"
      },
      {
        "type": "function_call_output",
        "call_id": "call_plan_001",
        "output": "{\"ok\":true,\"data\":{\"plan_code\":\"team\",\"display_name\":\"Team\",\"billing_mode\":\"usage_based\",\"included_requests\":10000,\"overage_allowed\":true,\"source_version\":\"fixture-2026-09-01\"}}"
      }
    ],
    "tools": [
      {
        "type": "function",
        "name": "get_plan_info",
        "description": "प्लान कोड से निश्चित परीक्षण डेटा प्राप्त करें",
        "strict": true,
        "parameters": {
          "type": "object",
          "properties": {
            "plan_code": {
              "type": "string",
              "enum": ["team"]
            }
          },
          "required": ["plan_code"],
          "additionalProperties": false
        }
      }
    ],
    "store": false
  }'

एक representative final output Item:

{
  "id": "resp_plan_002",
  "object": "response",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "Team प्लान शामिल सीमा से अधिक उपयोग पर usage-based billing को समर्थन देता है। परीक्षण डेटा में 10,000 requests शामिल हैं और overage_allowed true है।"
        }
      ]
    }
  ]
}

यदि server-side state continuation चुनते हैं, तो पहली response को store होने दें और दूसरी request में यह रूप उपयोग करें:

{
  "model": "YOUR_OPENAI_MODEL",
  "previous_response_id": "resp_plan_001",
  "input": [
    {
      "type": "function_call_output",
      "call_id": "call_plan_001",
      "output": "{\"ok\":true,\"data\":{\"plan_code\":\"team\",\"overage_allowed\":true}}"
    }
  ]
}

previous_response_id उसी upstream service का होता है जिसने response बनाया। इसे किसी अन्य provider को continuation के लिए नहीं दिया जा सकता। इससे historical input मुफ्त भी नहीं होता: OpenAI की वर्तमान documentation के अनुसार chain में पहले के input token अब भी input के रूप में bill होते हैं।

यदि response में reasoning Item हो, तो stateless replay में documentation के अनुसार संबंधित Item को भी बचाना होगा। “Uniform format” बनाने के लिए उसे हटाकर reasoning context को equivalent कहना सही नहीं है।

Anthropic Messages: पूरा tool round trip

Messages tool call को assistant content के tool_use block में रखता है और result को अगली user message के tool_result block में लौटाता है। Tool arguments पहले से object होते हैं, parse की प्रतीक्षा करती JSON string नहीं।

पहली request: Claude से tool_use लेना

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "YOUR_CLAUDE_MODEL",
    "max_tokens": 512,
    "system": "आप प्लान संबंधी सहायक हैं। केवल टूल से लौटे डेटा के आधार पर उत्तर दें; अनुमान न लगाएँ।",
    "messages": [
      {
        "role": "user",
        "content": "team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।"
      }
    ],
    "tools": [
      {
        "name": "get_plan_info",
        "description": "प्लान कोड से निश्चित परीक्षण डेटा प्राप्त करें",
        "strict": true,
        "input_schema": {
          "type": "object",
          "properties": {
            "plan_code": {
              "type": "string",
              "enum": ["team"]
            }
          },
          "required": ["plan_code"],
          "additionalProperties": false
        }
      }
    ],
    "tool_choice": {
      "type": "tool",
      "name": "get_plan_info"
    }
  }'

किसी विशेष tool को force करना selected model और configuration के support पर निर्भर है। Target model support न करे तो auto उपयोग करें और application में जाँचें कि वास्तव में tool call लौटा है।

एक representative response:

{
  "id": "msg_plan_001",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "tool_use",
      "id": "toolu_plan_001",
      "name": "get_plan_info",
      "input": {
        "plan_code": "team"
      }
    }
  ],
  "stop_reason": "tool_use"
}

input object को सीधे tool executor में दिया जा सकता है:

tool_result = execute_tool(
    "get_plan_info",
    {"plan_code": "team"},
)

दूसरी request: tool_result को तुरंत अगली user message में रखना

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "YOUR_CLAUDE_MODEL",
    "max_tokens": 512,
    "system": "आप प्लान संबंधी सहायक हैं। केवल टूल से लौटे डेटा के आधार पर उत्तर दें; अनुमान न लगाएँ।",
    "messages": [
      {
        "role": "user",
        "content": "team प्लान देखें और बताएँ कि शामिल सीमा से अधिक उपयोग पर usage-based billing उपलब्ध है या नहीं।"
      },
      {
        "role": "assistant",
        "content": [
          {
            "type": "tool_use",
            "id": "toolu_plan_001",
            "name": "get_plan_info",
            "input": {
              "plan_code": "team"
            }
          }
        ]
      },
      {
        "role": "user",
        "content": [
          {
            "type": "tool_result",
            "tool_use_id": "toolu_plan_001",
            "content": "{\"ok\":true,\"data\":{\"plan_code\":\"team\",\"display_name\":\"Team\",\"billing_mode\":\"usage_based\",\"included_requests\":10000,\"overage_allowed\":true,\"source_version\":\"fixture-2026-09-01\"}}"
          }
        ]
      }
    ],
    "tools": [
      {
        "name": "get_plan_info",
        "description": "प्लान कोड से निश्चित परीक्षण डेटा प्राप्त करें",
        "strict": true,
        "input_schema": {
          "type": "object",
          "properties": {
            "plan_code": {
              "type": "string",
              "enum": ["team"]
            }
          },
          "required": ["plan_code"],
          "additionalProperties": false
        }
      }
    ]
  }'

एक representative final response:

{
  "id": "msg_plan_002",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "Team प्लान शामिल सीमा से अधिक उपयोग पर usage-based billing को समर्थन देता है। परीक्षण डेटा में 10,000 requests शामिल हैं और overage_allowed true है।"
    }
  ],
  "stop_reason": "end_turn"
}

Messages में ordering स्पष्ट है: tool_result उस assistant message के तुरंत बाद होना चाहिए जिसमें matching tool_use है। यदि एक assistant turn कई client-side tool calls बनाता है, तो अगली user message में सभी संबंधित result blocks लौटाएँ और हर block को tool_use_id से जोड़ें। उसी user message में सामान्य text भी हो तो tool-result blocks को text से पहले रखें।

क्या सीधे map हो सकता है और कहाँ conversion में निश्चित रूप से हानि होगी

CapabilityConversion assessmentसही handling
सामान्य user textसामान्यतः सीधे map हो सकता हैकेवल दिखाई देने वाली strings नहीं, text, order और multimodal types सुरक्षित रखें
Basic function SchemaShape बदलकर map किया जा सकता हैfunction.parameters, Responses parameters और Messages input_schema के बीच convert करें, फिर supported JSON Schema subset को दोबारा validate करें
Tool argumentsType conversion आवश्यकदोनों OpenAI interfaces आम तौर पर JSON strings लौटाते हैं; Messages object लौटाता है। Business code से पहले normalize, parse और validate करें
Tool-call IDMeaning सुरक्षित रखें, namespace reuse न करेंInternal canonical call ID और original upstream ID दोनों रखें, फिर protocol-specific field से लौटाएँ
Parallel tool callsSupport संभव, पर array position से mapping नहींहर result को tool_call_id, call_id या tool_use_id से जोड़ें
system/developer instructionsLossy हो सकती हैंGlobal, conversation-stage और single-turn scope अलग करें; target protocol मूल scope व्यक्त न कर सके तो स्पष्ट downgrade या rejection करें
Final structured outputFields mechanically interchangeable नहींChat में response_format, Responses में text.format, Messages में output_config.format
Streamed tool argumentsProtocol-specific parser आवश्यकEvents और call ID के अनुसार fragments जमा करें, completion event के बाद JSON parse करें
Multi-turn server stateUniversal equivalent नहींprevious_response_id जैसे state IDs मूल upstream से जुड़े हैं; upstream बदलने पर visible context replay करें या sticky routing रखें
thinking/reasoning stateसामान्यतः lossless conversion संभव नहींProtocol-required opaque Items, thinking blocks, signatures या encrypted content को जस का तस रखें; स्वयं न बनाएँ
Hosted toolsअक्सर direct equivalent नहींweb search, file search, computer use, server tools जैसी capabilities और fallback अलग-अलग घोषित करें
Multiple candidatesEquivalent न हो सकता हैChat Completions के n को Responses में सीधे map मानने के बजाय application-level multiple requests या product behavior change करें

इसलिए gateway का सबसे भरोसेमंद internal abstraction एक विशाल object नहीं है जिसमें सारे fields मिला दिए जाएँ। Messages, instruction scope, tool definitions, tool calls, tool results, state handles, stream events और opaque native state को अलग-अलग model करें। कोई capability व्यक्त न हो सके तो field silently delete करने के बजाय स्पष्ट “unsupported” या “lossy conversion” status लौटाएँ।

strict, response_format और text.format अलग समस्याएँ हल करते हैं

Migration में सामान्य भ्रम है कि “tool arguments valid हैं” और “final answer specified JSON shape में है” एक ही feature हैं।

लक्ष्यChat CompletionsResponsesAnthropic Messages
Tool-call arguments सीमित करनाtools[].function.stricttools[].stricttools[].strict
Model का final output सीमित करनाresponse_formattext.formatoutput_config.format

Tool का strict नियंत्रित करता है कि मॉडल function को कैसे call करे। Final structured output नियंत्रित करता है कि user को क्या content मिले। Agent को दोनों साथ चाहिए हो सकते हैं: पहले strict arguments से tool call, फिर fixed JSON Schema में final result।

OpenAI की वर्तमान documentation में एक आसानी से छूटने वाला default अंतर भी है:

  • Chat Completions में function calls default रूप से non-strict होते हैं।
  • Responses में strict न देने पर service Schema को strict mode में normalize करने की कोशिश करती है। Incompatible होने पर non-strict mode पर लौट सकती है और parsed tool definition में strict: false दिखा सकती है।

Intent स्पष्ट रखने और interface-specific defaults पर निर्भरता से बचने के लिए production requests में strict: true या strict: false स्पष्ट रूप से set करें। Strict Schema को संबंधित requirements भी पूरी करनी चाहिए, जैसे extra object properties रोकना और सभी required fields देना।

और भी महत्वपूर्ण: Compatibility layer field स्वीकार कर सकती है, पर constraint लागू नहीं कर सकती। Anthropic की official OpenAI SDK compatibility documentation बताती है कि उस विशेष layer में function strict, response_format, reasoning_effort जैसे fields ignore होते हैं और अधिकतर unsupported fields error भी नहीं देते। इसलिए request 200 दे सकती है, जबकि Schema या reasoning setting वास्तव में लागू न हुई हो।

इसका अर्थ यह नहीं कि native Anthropic Messages में equivalent capabilities नहीं हैं। Native Messages strict tool input support करता है और final JSON के लिए output_config.format उपयोग करता है। Troubleshooting की पहली जाँच होनी चाहिए: call native Messages पर जा रही है या OpenAI-compatible layer पर?

system, developer और instruction scope केवल strings जोड़कर सुरक्षित नहीं रहते

OpenAI-style interfaces message history में अलग roles देते हैं, और Responses में instructions भी है। Anthropic Messages लंबे समय से top-level system उपयोग करता रहा है। सितंबर 2026 तक कुछ current models mid-conversation role: "system" भी support करते हैं, पर सभी नहीं; placement और tool-ordering constraints भी लागू हैं।

साथ ही Anthropic की OpenAI SDK compatibility layer conversation के system/developer messages जमा करती है, newline से जोड़ती है और एक शुरुआती system prompt में promote कर देती है। Request usable हो जाती है, पर मूल timing और scope बदल जाते हैं। केवल आठवें turn से लागू होने वाली developer instruction शुरू में पहुँचकर पहले सात turns की semantics भी बदल सकती है।

सुरक्षित adapter पहले application में तीन scopes अलग करता है:

  • Global instructions: पूरी conversation पर लागू।
  • Conversation-stage instructions: किसी निश्चित turn से आगे लागू।
  • Single-turn instructions: केवल current task को नियंत्रित करती हैं।

Target protocol वही scope व्यक्त कर सके तभी mapping करें। न कर सके तो स्पष्ट strategy चुनें: capability-supporting model पर fixed routing, instruction downgrade करके difference record करना, या migration reject करना। Silent concatenation कम code लेती है, पर “request सफल, behavior बदल गया” वाली समस्या सबसे अधिक बनाती है।

Streaming को text token जोड़ने के बजाय state machine की तरह parse करें

तीनों interfaces streaming support करती हैं, पर events equivalent नहीं हैं:

  • Chat Completions सामान्यतः choices[].delta से text और tool_calls fragments जमा करता है।
  • Responses response.output_text.delta, response.function_call_arguments.delta, response.function_call_arguments.done, response.completed और error जैसे typed events उपयोग करता है।
  • Messages message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop उपयोग करता है; tool arguments input_json_delta.partial_json से fragments में आते हैं।

Tool arguments इस तरह विभाजित हो सकते हैं:

{"plan_
code":"te
am"}

इनमें से कोई fragment अकेले valid JSON नहीं है। Call ID या content-block index के अनुसार जमा करें और arguments completion event मिलने के बाद parse करें:

from __future__ import annotations

import json
from collections import defaultdict
from typing import Any


class ToolArgumentAssembler:
    def __init__(self) -> None:
        self._buffers: dict[str, list[str]] = defaultdict(list)

    def add_delta(self, call_id: str, fragment: str) -> None:
        self._buffers[call_id].append(fragment)

    def finish(self, call_id: str) -> dict[str, Any]:
        if call_id not in self._buffers:
            raise KeyError(f"अज्ञात call_id: {call_id}")

        raw = "".join(self._buffers.pop(call_id))
        value = json.loads(raw)
        if not isinstance(value, dict):
            raise TypeError("टूल arguments को object के रूप में decode होना चाहिए")
        return value

    def discard(self, call_id: str) -> None:
        self._buffers.pop(call_id, None)

Adapter को स्पष्ट terminal state भी record करनी चाहिए:

created -> receiving -> completed
                   \-> failed
                   \-> disconnected

disconnected, completed नहीं है। Anthropic Messages HTTP connection सफल होने के बाद भी stream के अंदर event: error दे सकता है; Responses में भी अलग error events हैं। केवल शुरुआती HTTP status देखना या connection close को natural completion मानना tool arguments या final answer काट सकता है।

Event parser को unknown event types भी tolerate करने चाहिए: current capability पर असर न डालने वाले events को log करके skip करें, हर नए server event पर पूरा client crash न करें।

Multi-turn state और reasoning state गढ़े नहीं जा सकते

Chat Completions और पारंपरिक Messages flow सामान्यतः application द्वारा history replay पर निर्भर हैं। Responses previous_response_id या Conversations से server-side state भी रख सकता है। इनका “previous turn” एक-दूसरे के बराबर नहीं है।

Gateway को state ID मिलने पर केवल तीन सही strategies हैं:

  1. Sticky routing: बाद की requests उसी upstream को जाएँ जिसने state बनाया।
  2. Complete replay: legally replay होने वाले messages, tool calls, tool results और native state elements दोबारा भेजें।
  3. Explicit rejection: target upstream continue न कर सके तो diagnosable error लौटाएँ और client को conversation फिर शुरू करने दें।

OpenAI का previous_response_id Anthropic को न भेजें और gateway की internal conversation ID को किसी अन्य provider द्वारा समझे जाने वाला state handle न मानें।

Reasoning state भी field rename करने से नहीं बदलता:

  • Stateless या कुछ data-retention configurations में Responses encrypted reasoning Items लौटा सकता है जिन्हें बाद की request में replay करना पड़ता है।
  • Anthropic thinking workflows में thinking blocks, signatures या अन्य opaque state हो सकती है। Tools और multi-turn conversation में इसे native documentation के अनुसार रखना चाहिए।
  • सितंबर 2026 तक Anthropic का manual thinking.type: "enabled" और budget_tokens संयोजन 4.6-generation models में deprecated है और 4.7 तथा बाद की generations में reject होता है; नए models adaptive thinking और संबंधित effort control उपयोग करते हैं।

इसलिए OpenAI reasoning_effort को Anthropic budget_tokens के बराबर बताने वाला स्थायी नियम नहीं बनाया जा सकता। सही capability description में target model, current thinking mode और unsupported होने पर fallback behavior शामिल होना चाहिए।

Parallel tool calls: array position नहीं, ID से जोड़ें

Model एक turn में कई tools माँग सकता है। Execution time अलग होने से results का order calls के order से अलग हो सकता है। Adapter को इस तरह का relation रखना चाहिए:

canonical_call_id
  -> provider
  -> provider_call_id
  -> tool_name
  -> validated_arguments
  -> execution_status
  -> result

Results लौटाते समय:

  • Chat Completions हर result के लिए role: "tool" message बनाता है और matching tool_call_id देता है।
  • Responses हर result के लिए function_call_output Item बनाता है और matching call_id देता है।
  • Messages तुरंत अगली user turn में matching tool_result blocks रखता है और हर block में संबंधित tool_use_id देता है।

Migration testing में पहले parallel_tool_calls: false रखें, single-tool path पूरा चलाएँ और फिर parallelism enable करें। Production में email भेजने, charge करने या resource बनाने जैसे side-effect tools को idempotency keys भी चाहिए। Network retry, stream disconnect या upstream replay उसी semantic call को फिर पहुँचा सकते हैं; model-generated text से यह तय न करें कि call पहले execute हुई थी या नहीं।

Compatibility validate करने के लिए HTTP 200 पर्याप्त क्यों नहीं

एक उपयोगी migration test को कम से कम ये paths cover करने चाहिए:

TestPassing criterion
सामान्य textContent पढ़ने योग्य हो और system/developer scope expected behavior दे
Single tool callTool name, arguments, call ID, result और final answer पूरा round trip बनाएँ
Parallel tool callsहर result ID से सही जुड़ा हो, cross-link या missing call न हो
Streamed tool argumentsFragments पूरी तरह जुड़ें और completion के बाद JSON parse हो
Strict tool SchemaInvalid fields/types expected तरीके से reject हों या explicit downgrade हो
Final structured outputFinal answer specified Schema को पूरा करे, केवल “JSON जैसा” न दिखे
Tool execution errorModel को structured error मिले, infinite repeat या fake success न हो
Multi-turn continuationदूसरा turn पहले turn के facts refer कर सके और state-switch rules स्पष्ट हों
reasoning/thinkingDeclared mode काम करे और native state delete या fabricate न हो
In-stream errors/disconnectsClient completion, failure और connection interruption अलग कर सके
Controlled API errorError type, request ID और retry policy diagnosable रहें

Fixed input और fixed tool fixture उपयोग करें तथा हर protocol के लिए अलग record करें:

  • final business result equivalent है या नहीं;
  • tool call और result handoff पूरा है या नहीं;
  • P50 और P95 latency;
  • input, output और cache-related usage;
  • error type, request ID और terminal state;
  • कौन-सी capabilities explicitly downgraded हुईं।

API Key, पूरा sensitive prompt या user private output log न करें। Error logs में कम से कम HTTP status, upstream error type/code, short message, request ID, endpoint, protocol, Model ID और stream terminal state रखें। अन्यथा model_not_found, missing permission और incompatible path सब एक undiagnosable 400 बन सकते हैं।

Symptom के अनुसार diagnosis: Agent कहाँ टूटा

Symptomसामान्य कारणजाँच और समाधान
200 लौटता है पर model कभी tool call नहीं करताTool definition नहीं भेजी, tool_choice ignore, model tools support नहीं करता, prompt पर्याप्त नहींFinal outbound request print करें; target model और compatibility layer जाँचें; test में केवल एक read-only tool दें और use force या explicitly request करें
Model tool call देता है पर application execute नहीं करतीपुराना field पढ़ रही है, जैसे केवल message.contentProtocol के अनुसार tool_calls, function_call Item या tool_use block पढ़ें
Tool arguments JSON parse failStreaming fragment को complete JSON माना या object को दोबारा string की तरह parse कियाArguments completion event का इंतजार करें; पहले देखें value string है या object
strict के बाद भी extra fieldsCompatibility layer silently ignore, Schema strict rules पूरी नहीं करती, native endpoint नहींFinal endpoint और documentation जाँचें; strict explicit set करें; जानबूझकर Schema violation वाला regression test जोड़ें
दूसरी request tool result missing कहती हैCall ID mismatch या पहला assistant/tool Item save नहीं हुआUpstream call और ID को unchanged रखें; protocol-required position पर तुरंत result लौटाएँ
Messages tool_use ids ... without tool_result देता हैtool_result call के तुरंत बाद नहीं या पहले सामान्य text हैसभी matching tool_result अगली user message में और optional text से पहले रखें
Streaming रुकती है या आधे arguments मिलते हैंClient केवल text end marker की प्रतीक्षा करता है, argument/error terminal states नहींहर protocol के लिए event state machine बनाएँ और completed, failed, error, disconnected अलग करें
दूसरा turn पहला भूलता हैHistory, tool calls या result Items missing; या previous_response_id दूसरे upstream कापूरा visible context replay करें या sticky routing रखें; providers के बीच state IDs न भेजें
Interface बदलने पर system instruction बहुत जल्दी लागूCompatibility layer ने mid-conversation system/developer को शुरुआत में promote कियाInstruction scope model करें; lossless mapping न हो तो explicit downgrade या native protocol fixed रखें
Tool दो बार executeRequest retry, disconnect के बाद replay, idempotency control नहींTesting में read-only tools; production side-effect tools के लिए canonical call ID से idempotency key बनाएं
Final content JSON है पर fields कभी missingPrompt केवल “JSON लौटाओ” कहता है, structured output enabled नहींसंबंधित interface का response_format, text.format या output_config.format उपयोग करें, फिर application में validate करें

अधिक सुरक्षित migration sequence

  1. Client वास्तव में कौन-सा protocol भेजता है, पता करें। Model name से अनुमान न लगाएँ। Full endpoint, SDK method, top-level request fields और stream event types record करें।
  2. Preserve होने वाले behavior की सूची बनाएँ। कम से कम tools, parallel calls, strict Schema, final structured output, multi-turn state, streaming और thinking/reasoning।
  3. पहले native protocol चुनें। Native Messages या Responses से संभव capability के लिए अतिरिक्त compatibility layer से बचें।
  4. Conversion capability matrix बनाएं। हर capability को fully supported, lossy support या unsupported चिन्हित करें और caller को result दिखे।
  5. Side-effect-free fixture से पूरा two-request loop चलाएँ। केवल “एक request से text मिला” पर न रुकें; tool execute करें और result वापस दें।
  6. फिर parallel, streaming और error paths test करें। Normal path pass होने के बाद ही side-effect tools और real traffic खोलें।
  7. Traffic धीरे बढ़ाकर metrics compare करें। केवल HTTP success rate नहीं; correctness, latency, usage, errors और duplicate tool execution भी देखें।

BetterToken में सही entry point चुनना

BetterToken अलग clients के लिए अलग connection paths देता है। Protocol वही होगा जो client वास्तव में wire contract के रूप में उपयोग करता है:

  • Chat Completions: पूरा request URL https://www.bettertoken.ai/v1/chat/completions है। Path अपने आप जोड़ने वाले SDK या tools में Base URL सामान्यतः https://www.bettertoken.ai/v1 रखें। Chat Completions API reference देखें।
  • Codex / Responses: वर्तमान Codex documentation base_url = "https://www.bettertoken.ai/v1" और wire_api = "responses" उपयोग करती है; Codex स्वयं /responses जोड़ता है। Codex setup guide देखें।
  • Claude Code / Messages: वर्तमान documentation ANTHROPIC_BASE_URL=https://bettertoken.ai उपयोग करती है और Base URL के बाद /v1 नहीं जोड़ती; client /v1/messages जोड़ता है। Claude Code setup guide देखें।

एक ही Dashboard, API Key या model name तीन protocols को एक format नहीं बनाता। Existing tool में वही protocol चुनें जिसकी उसे अपेक्षा है। अपना Agent बनाते समय इस लेख के complete round trips और acceptance matrix से आवश्यक capabilities verify करें।

सामान्य प्रश्न

क्या OpenAI-compatible का अर्थ OpenAI API की पूरी copy है?

नहीं। सामान्यतः इसका अर्थ है कि कुछ endpoints और data structures को OpenAI-style clients call कर सकते हैं। Models, parameters, stream events, tools, structured output, hosted tools और error semantics को अलग-अलग verify करना पड़ता है।

क्या केवल Base URL और API Key बदलना पर्याप्त है?

कभी-कभी, यदि सरल text requests पहले से वही wire contract उपयोग करती हों। Tool Agent में tool definitions, दूसरी request का result handoff, stream events, strict Schema, state और errors फिर भी validate करने होंगे। Client Responses अपेक्षा करे तो केवल /chat/completions पर्याप्त नहीं; Messages अपेक्षा करे तो OpenAI-style endpoint अपने आप adapt नहीं होगा।

क्या एक universal adapter तीनों protocols convert कर सकता है?

यह सामान्य text और function-tool loop का कुछ भाग संभाल सकता है, पर full lossless support का दावा नहीं करना चाहिए। Provider-managed state, hosted tools, opaque thinking/reasoning state, कुछ system scopes और model-specific capabilities का universal equivalent अक्सर नहीं होता। Adapter को capability matrix और downgrade information दिखानी चाहिए।

Unit tests pass होते हैं पर वास्तविक Agent fail क्यों करता है?

बहुत से tests केवल पहली model response mock करते हैं। वे दूसरी request का tool-result handoff, parallel calls, streaming fragments या state continuation verify नहीं करते। Test को “user request → model tool call → application execution → tool result handoff → final response” तक बढ़ाएँ ताकि वास्तविक protocol problems सामने आएँ।

Migration में पहले Chat Completions या Responses बदलना चाहिए?

Stable Chat Completions application चलती रह सकती है और business value के अनुसार capability-by-capability migrate कर सकती है। नया OpenAI Agent, या typed Items, hosted tools अथवा Responses state capabilities की स्पष्ट जरूरत वाला Agent, सीधे Responses पर शुरू करना बेहतर है। निर्णय capabilities और migration cost पर होना चाहिए, interface name के नए या पुराने लगने पर नहीं।

अपना LLM वर्कफ़्लो बेहतर बनाना चाहते हैं?

एक API से मॉडल जोड़ें, कुंजियाँ प्रबंधित करें और AI खर्च नियंत्रित करें।

मुफ़्त शुरू करें