आमंत्रित करें और कमाएँ

आमंत्रण पुरस्कार कैसे काम करते हैं

अपना आमंत्रण लिंक साझा करें। मित्र इसके माध्यम से पंजीकरण करके टॉप-अप करता है तो उसके बाद के टॉप-अप पर आपको दिखाया गया पुरस्कार मिलेगा।

API लागत calculator: tokens, cache और request volume

Token rates और request volume से forecast बनाना, cache hit को अलग रखना और accepted workflow की वास्तविक लागत जाँचना।

विषय-सूची

यह interactive web form नहीं है। यह calculator एक transparent formula और local Python script है: input, output, cache write, cache read, calls और current rates आप स्वयं भरते हैं। Unknown value को चुपचाप zero न बनाएं। सभी rates एक currency में प्रति 1,000,000 tokens रखें।

calculator को कौन-से data चाहिए

input_tokens_per_call
output_tokens_per_call
cache_write_tokens_per_miss
cache_read_tokens_per_hit
calls
cache_hit_rate
prices_per_1m_tokens

क्या आप forecast को real call से मिलाना चाहते हैं? BetterToken account और API Key बनाएँ, pricing page से current rates लें और एक controlled request चलाएँ। फिर Dashboard में model, status, input, output, cache Token और charge मिलाएँ।

universal formula

I  — regular input tokens
O  — output tokens
W  — cache write / creation tokens
R  — cache read / cached tokens
Pi — input price per 1,000,000 tokens
Po — output price per 1,000,000 tokens
Pw — cache write price per 1,000,000 tokens
Pr — cache read price per 1,000,000 tokens
C = I / 1_000_000 × Pi
  + O / 1_000_000 × Po
  + W / 1_000_000 × Pw
  + R / 1_000_000 × Pr
  + Cextra

Cextra अलग billed units के लिए है। कोई category unknown हो तो उसे unknown रखें, zero मानकर false precision न दें।

copyable Python calculator

from decimal import Decimal, InvalidOperation

MILLION = Decimal("1000000")

def read_decimal(label: str, *, allow_empty: bool = False) -> Decimal:
    raw = input(label).strip().replace(",", ".")
    if allow_empty and raw == "":
        return Decimal("0")
    try:
        value = Decimal(raw)
    except InvalidOperation as exc:
        raise SystemExit(f"Invalid number for {label!r}") from exc
    if value < 0:
        raise SystemExit(f"Negative value is not allowed for {label!r}")
    return value

input_tokens = read_decimal("Input tokens per call: ")
output_tokens = read_decimal("Output tokens per call: ")
cache_write_tokens = read_decimal("Cache write tokens per call: ")
cache_read_tokens = read_decimal("Cache read tokens per call: ")
calls = read_decimal("Number of calls: ")

price_input = read_decimal("Input price per 1M tokens: ")
price_output = read_decimal("Output price per 1M tokens: ")
price_cache_write = read_decimal("Cache write price per 1M tokens: ")
price_cache_read = read_decimal("Cache read price per 1M tokens: ")
extra_per_call = read_decimal("Extra cost per call (empty = 0): ", allow_empty=True)

per_call = (
    input_tokens / MILLION * price_input
    + output_tokens / MILLION * price_output
    + cache_write_tokens / MILLION * price_cache_write
    + cache_read_tokens / MILLION * price_cache_read
    + extra_per_call
)
total = per_call * calls

print(f"Cost per call: {per_call:.8f}")
print(f"Total cost:    {total:.8f}")

cache hit rate और तीन scenarios

Hits और misses अलग करें: Nhits = N × h, Nmiss = N - Nhits, और Ctotal = Nhits × Chit + Nmiss × Cmiss + Cextra_total। Planning में hits को नीचे और misses को ऊपर round करें। Base scenario में median usage, favorable scenario में stable prefix/high hit rate, और worst case में cache misses, long output, bounded retry तथा separate tools रखें।

estimate को actual से मिलाएँ

एक safe test task चलाएँ, actual API calls को model और usage category के अनुसार group करें, formula apply करें और Dashboard से compare करें। समय, Model ID, input/output, cache category, retries, actual charge, currency और price date check करें।

सामान्य input rate से पहले usage-field guide से total input context में शामिल cache read अलग करें, वरना double count हो सकता है। Formula future prices, dynamic routing या unknown tool units predict नहीं करता।

FAQ

Cache write/read को zero केवल तभी रखें जब endpoint ने वास्तव में cache नहीं लगाया हो। अलग models और tasks को अलग rows में गिनें; actual charge अधिक हो तो rates, output, retries, cache misses और extra tools जाँचें।

अपना LLM वर्कफ़्लो बेहतर बनाना चाहते हैं?

एक API से मॉडल जोड़ें, कुंजियाँ प्रबंधित करें और AI खर्च नियंत्रित करें।

मुफ़्त शुरू करें