Pricing
Per million tokens, one balance across every live model. No subscription, no minimum spend, and credit you buy does not expire. Every number on this page renders from the same published facts file the console bills from. Weighing this against self-hosting or a rented GPU?
rate quote · GLM-5.3-Flashpriced 2026-09-09
- input
- $0.15
- output
- $0.50
- cached input
- $0.03
- context
- 1M tokens
- output speed
- up to 179 tok/s
per 1M tokens, USD: a dollar buys 6.7M input tokens. Speed measured 2026-09-13, sealed protocol.
| Model | Input | Cached input | Output | Context | Output speed |
|---|---|---|---|---|---|
| GLM-5.3-Flash chat and tools | $0.15 | $0.03 | $0.50 | 1M | up to 179 tok/s |
| DeepSeek-V4.1-Flash chat and tools | $0.30 | $0.006 | $1.20 | 1M | up to 101 tok/s |
| Qwen3-Embedding-8B embeddings | $0.02 | $0.00 | no output tokens | 16K | not measured |
| Qwen3-Reranker-8B reranking | $0.05 | $0.00 | no output tokens | 16K | not measured |
- Input
- Every prompt token the model had to read for the first time. Billed as
usage.prompt_tokens. - Cached input
- Prompt tokens served from a cached prefix, billed below ordinary input. Applied automatically; cache writes carry no separate fee. Reported as
usage.prompt_tokens_details.cached_tokens. Exception, GLM-5.3-Flash: Prompt caching restores retained prefixes and bills only the tokens reported as cached_tokens. In qualification, repeating a 902,579-token prompt reused 902,560 tokens; warm repeat TTFT was 1.819 to 3.972 seconds. Cold prompts still need processing. Residency is bounded and shared; a later request can return cold. Exception, DeepSeek-V4.1-Flash: Prompt caching restores retained prefixes and bills only the tokens reported as cached_tokens. In qualification, a 32.7k-token prompt (32,674-32,758 tokens) reported 32,512 cached tokens on 9 of 9 repeats. Cold prompts still need processing. Residency is bounded and shared; a later request can return cold. - Output
- Tokens the model generated, including reasoning tokens when thinking is on. Billed as
usage.completion_tokens. - Per request
- $0.00. No request fee, no per-key fee, no minimum spend.
- Context
- The full native window at the rates above. Long prompts bill more tokens, never a higher rate: there is no long-context tier and no separate charge for the window.
- Output speed
- A measured ceiling on the serving build, input to output: the fastest measured shape, so ordinary requests land below it. First-token figures are published per model, with reasoning off, on the model pages. The measurement protocol.
- Priced
- The date the current rate was set. Sort by it to see what is newest.
* GLM-5.3-Flash is served with DFlash 2 (inco.ai) as its speculative drafter used by this model, under written permission from its authors.
Measurement receipts
GLM-5.3-Flash: up to 179 tok/s, measured 2026-09-13, build verified 2026-09-13.
DeepSeek-V4.1-Flash: up to 101 tok/s, measured 2026-09-16, build verified 2026-09-16.
Single stream, measured on the live endpoint: observations, not service guarantees. The full protocol, vantage and build behind each figure are on the measurements page.
One prepaid balance and one key cover every model; each request bills at the rate of the model it named. Every live model is served at the same base URL; the model id in the request picks the rate. The quickstart shows the request shape. Every response reports billed token counts split into input, cached input and output, in the standard usage field, in streaming responses too. The same numbers appear in your console.
| Model | Streaming | Whole file | Language | Endpoint |
|---|---|---|---|---|
Whisper Large v3 Hebrewivrit-ai/whisper-large-v3-he | $0.18 | $0.10 | Hebrew | /v1/realtime streaming/v1/audio/transcriptions whole file |
Qwen3 ASR Englishqwen/qwen3-asr-en | $0.15 | $0.04 | English | /v1/realtime streaming/v1/audio/transcriptions whole file |
One prepaid balance and one key cover these too, and they answer the same base URL as every other model. A transcription response reports the billed audio seconds in itsusage field, the same place a chat response reports billed tokens, and the same number appears in your console. Request shapes are in the API documentation.
Running this at company scale?
Talk to us directly and we set up an arrangement for your workload: volume, limits, billing, terms. You reach Avi Fenesh, founder, not a sales queue.
Need the model tuned to your data, or run on hardware you control? The lab's services.
Automatic top-up
Off by default. You set the trigger, the amount, and a monthly ceiling.
- Default trigger
- $5 remaining
- Default purchase
- $25 of credit
- Monthly ceiling
- You set it, default $200. We stop at the ceiling and email you rather than charging past it.
- If a charge fails
- Auto top-up switches itself off after two failures and emails you.
- Turning it off
- One click in the console, effective immediately.
Starting out
New accounts pay per use, no minimum, no subscription. Credit you buy is metered at the rates above and does not expire.
- Rate limit
- Limits exist per account for abuse control and are raised on request.
- Data use
- Purchased-credit traffic is not used for training or evaluation by default; an account setting can opt in for the rebate described below. The full posture.
- Train on my prompts, 5% back
- One switch in the console: agree that your prompts may be used to improve our models and 5% of what you spend comes back as credit, added a day at a time for each full day the switch is on. The setting is off by default and can be changed at any time. Rebate is calculated from eligible metered traffic.
The table is the whole price.
No request fee, no per-key fee, no long-context tier, no contract.
Rates current as of 2026-09-16. Live status: status.tiyuvta.ai.