208 callable models — 105 text across 12 vendors, 25 image, 70 video, generated from the pricing registry
The platform currently offers 208 callable models: 105 text chat models across 12 vendors, plus 25 image and 70 video generators, and 8 music / voice / avatar SKUs (last section on this page). Text pricing is in credits per 1M tokens; images bill per image and video per second.
This page is generated from the same pricing registry the gateway bills against, so every id here is callable and every price is the one you are charged.
⚠️ The 208 here and the 213 on the model catalogue count different things — this is not drift: the plaza additionally shows not-yet-launched placeholder cards (badged "coming soon"), while this page lists only ids you can call today. A few API-only SKUs appear on neither. Live data: GET https://oemoemapi.dflop.top/api/v1/models/public, or the model catalogue.
Cached pricing needs no opt-in: when upstream usage reports cached_tokens, the gateway settles that portion of the input at the cached rate automatically — clients don't (and can't) enable it explicitly. Models without a cached figure bill all input at the Input price.
Only Anthropic Claude charges for cache writes. The first time Claude writes a prefix into the cache, those tokens bill at Input × 1.25 (cache_creation_input_tokens); later hits bill at the cached rate above (Input × 0.1). The cache lives about 5 minutes, refreshed on every hit — if no follow-up turn reuses the prefix within that window, the write premium buys nothing. Every other vendor (OpenAI / xAI / Alibaba / Zhipu / Moonshot / DeepSeek / MiniMax / Tencent / Volcengine) only has a hit discount and does not bill writes separately.
Limited-time promotion: the whole Anthropic Claude line is charged at 40% off (list price × 0.6) and the whole OpenAI GPT text line at 70% off (× 0.3), with cached rates discounted equally. The tables show list prices; actual billing always uses the discounted rate, for both API-key calls and in-platform conversations. Discounted models carry an "X% OFF" badge on the model catalogue.
Long-context tier: for grok-4.5 / grok-4.6, a prompt of ≥200K tokens bills input, cache-hit and output all at 2× the table price. That is xAI's whole-turn doubling for very long requests, passed straight through (verified to the cent on 2026-08-23). The trigger is the total prompt size of the turn, independent of how much was generated. No other model has this tier.
Server-side tools bill per call: when a grok model runs web search / tools, the upstream charges 2.022 points per tool call on top of tokens. The count comes from the upstream's usage.num_server_side_tools_used and is written to the server_tool_calls column of every usage-log row, so you can re-derive your bill. Turns that use no tools incur no such charge.
gpt-image-2, gpt-image-2.5-flare and gpt-image-2.5-sunburst all accept POST /v1/images/edits for reference-image edits (sketch-to-image works the same way). The upstream ignores size — put the aspect ratio in the prompt; latency is 30–215s.
The two Midjourney entries are priced per image, but one request always produces 4 of them (a 2×2 grid) and n does not change that — a single request costs 161.76 (v8.1) / 129.41 (v7) credits. size only sets the aspect ratio, never exact pixels; the output tier goes at the end of the prompt (--sd / --hd). Integration details in Image / video APIs.
Since 2026-07-08 the three nano entries are the only way in: older ids (gemini-2.5-flash-image, gemini-3-pro-image-preview, gemini-3.1-flash-image(-preview), tvod-nano-*) keep working as aliases onto the canonical entries above.
Billed per second as an async task through POST /v1/videos/generations. Where tiers are listed (resolution, or operation for the subtitle SKU) you are billed at the tier you request. See Image / video APIs.
480p 27.72144 / 720p 59.616 / 1080p 148.716 per second; token rate by delivery tier: 480p 2760/1M (no video input) 1680/1M (with video input), 720p 2760/1M (no video input) 1680/1M (with video input), 1080p 3060/1M (no video input) 1860/1M (with video input)
doubao-seedance-2.0-fast
Seedance 2.0 Fast
480p 22.29768 / 720p 47.952 / 1080p 107.892 per second; token rate by delivery tier: 480p 2220/1M (no video input) 1320/1M (with video input), 720p 2220/1M (no video input) 1320/1M (with video input), 1080p 2220/1M (no video input) 1320/1M (with video input)
doubao-seedance-2.0-fast-lite
Seedance 2.0 Fast Lite
720p 47.952 / 1080p 107.892 per second; billed by token = token rate x tokens + upscale rate x delivered seconds. Token rate 720p 4525.29/1M (no video input) 2690.71/1M (with video input), 1080p 4763.52/1M (no video input) 2832.36/1M (with video input). Upscale add-on: 720p 2.5/s, 1080p 5/s (charged on delivered seconds only, never on the reference video)
doubao-seedance-2.0-lite
Seedance 2.0 Lite
720p 59.616 / 1080p 148.716 per second; billed by token = token rate x tokens + upscale rate x delivered seconds. Token rate 720p 5686.58/1M (no video input) 3461.4/1M (with video input), 1080p 6653.52/1M (no video input) 4049.97/1M (with video input). Upscale add-on: 720p 2.5/s, 1080p 5/s (charged on delivered seconds only, never on the reference video)
doubao-seedance-2.0-mini
Seedance 2.0 Mini
480p 13.86072 / 720p 29.808 per second; token rate by delivery tier: 480p 1380/1M (no video input) 840/1M (with video input), 720p 1380/1M (no video input) 840/1M (with video input)
doubao-seedance-2.0-mini-lite
Seedance 2.0 Mini Lite
720p 29.808 / 1080p 74.358 per second; billed by token = token rate x tokens + upscale rate x delivered seconds. Token rate 720p 2718.84/1M (no video input) 1654.95/1M (with video input), 1080p 3211.02/1M (no video input) 1954.53/1M (with video input). Upscale add-on: 720p 2.5/s, 1080p 5/s (charged on delivered seconds only, never on the reference video)
doubao-seedance-2.5
Seedance 2.5
480p 40.3515 / 720p 90.72 / 1080p 224.532 per second; token rate by delivery tier: 4200/1M (no video input), 2520/1M (with video input); the 1080p tier bills 4620/1M and 2760/1M
doubao-seedance-2.5-lite
Seedance 2.5 Lite
720p 42.8515 / 1080p 95.72 per second; billed by token = token rate x tokens + upscale rate x delivered seconds. Token rate 4200/1M (no video input), 2520/1M (with video input). Upscale add-on: 720p 2.5/s, 1080p 5/s (charged on delivered seconds only, never on the reference video)
Both Wan 3.0 cards share one capability set: 2-30s, 480P/720P/1080P, 30fps, aspect 16:9 / 4:3 / 1:1 / 3:4 / 9:16 / adaptive; text-to-video, image-to-video (first frame, first+last frame), and image / video / audio references. Prime is the accelerated tier (same quality, markedly faster) at 1.5x the standard rate. The default delivery is 1080P — pass resolution explicitly for a cheaper tier.
⚠️ On rounds that carry a reference video, the input video's seconds are billed too. This is the upstream's own basis: billed duration = input duration + output duration, and the rate above applies to both legs (upstream's price table heads that column "input and output unit price"). Reference videos are [1,15]s each and 15s total; we hold against that ceiling at submit time and refund the difference once upstream reports the real input duration. Reconcile against the top-level input_video_duration_sec in the response (duration_sec still means the delivered clip only, excluding the input). Text-to-video and image-to-video rounds are unaffected (image input is free upstream, and we don't charge for it either). See Media APIs - how usage relates to billing.
Different models support different client SDKs, and the compatibility matrix is authoritative on what is actually reachable. The catalog's supported_protocols field reflects only the native / primary path; many models are additionally reachable on other protocols through gateway translation (the response then carries an X-Protocol-Translation header). Native paths, counted from the registry that generated this page:
OpenAI Chat (/v1/chat/completions) — all 105 text models
For a model × protocol combination with no reachable channel, the gateway returns 503 no_channel_available (the model exists but has no channel on that protocol — not a 404); switch to a supported protocol.
134 older ids still resolve onto the canonical entries above, so integrations pinned to them keep working. An id that is neither in the tables above nor an alias returns model_not_found — the catalog endpoint is the source of truth.
The response looks like { "models": [...] }. Note that this endpoint returns every registered entry — including placeholder SKUs that aren't live yet and non-chat SKUs — so filter on the three fields below when consuming it from a script rather than looping over the whole list:
Field
Meaning
callable
false means a placeholder SKU (listed but not live); calling it returns model_not_found
endpoint_type
null means an ordinary chat model. images_generations, videos_generations, contents_generations_tasks and similar mean the model uses its own dedicated endpoint and cannot be sent to /v1/chat/completions — see Image / video / music APIs
supported_protocols
the client protocols available for this model (see above)
To take only the models you can send straight to /v1/chat/completions: