models

The roster.

One base URL; the model field picks the route. Each model's page carries three things kept deliberately separate: its maker's own card, the benchmarks its maker publishes, and the numbers we measured on the live endpoint. Deciding between this and running it yourself? Self-hosting, a rented GPU, or hosted, costed.

live routes6one API · one account balance

Generation

2 models · chat, tools, vision

Rail figures are ours, measured on the live endpoint: first token is the published figure at the fast edge, pinned to its build, never a median; where the model has a reasoning-off mode it is measured with reasoning off, otherwise it is the first streamed token with the default thinking on (each model page carries the current measurement record). GLM-5.3-Flash publishes no first-token figure. Cold cache, and the RTT between you and the edge is not part of the number; output speed is a single-stream ceiling. Dates, builds and full conditions are on each model’s page.

Speech to text

2 models · streaming & whole file · billed by the audio hour

No speed cell on purpose, and not because one is pending: what we measure on these models is how many real-time streams a card carries, which is a capacity figure we keep internal, not a customer claim. Billing is audio duration, silence included, with no minimum per request: an hour of audio costs the hourly rate whether it arrives as one file or as a live stream. Request shapes are in the API documentation.

Retrieval

2 models · embeddings & rerank · dedicated capacity

No speed cell on purpose. Since 2026-09-02 retrieval hasdedicated capacity: it is admitted up to the advertised concurrency without yielding to the generation lanes, and only beyond that concurrency does the lane cap shed with a retryable 429. That posture is the price basis. Billing is input tokens only.

Whose numbers are whose

the rule every page here follows

The maker’s numbers

Benchmark scores on a model’s page are its maker’s own published figures, quoted with sources and retrieval dates, not our measurements, and not comparable between makers, who each ran their own harness.

Our numbers, and the gates behind them

First-token and output-speed figures are ours, measured on the live endpoint: single stream, published reference figures (never medians), cold cache, conditions and build pins on each model’s page. Every build passes correctness gates before it takes traffic: what we serve is reconciled against the vendor’s published model, and a live gate covers all three wire formats plus a tool-call round trip. These are observations, not guarantees; run your own and use those instead.