Published rates
The published rates are effective August 29, 2026. Catalogue, availability and prices may change. An accepted Palace Ring Order or institutional agreement controls.
Service 01
Published usage rates for text, image, video, audio and retrieval models.
Public rate schedule
Rates last verified August 29, 2026. Model availability and pricing remain subject to the applicable Order.
72 listed services
USD per 1M tokens. A dash means no separate cached-input rate is listed. Context is maximum tokens where published.
| Model | Workload | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
| Kimi K3 | Chat, reasoning, coding agents, vision | 1,048,576 | $3.00 | $0.30 | $15.00 |
| Qwen3.8 2.4T A95B | Large-scale reasoning and vision | — | $2.50 | $0.50 | $6.25 |
| GLM-5.3 | Reasoning, agents and structured work | 1,000,000 | $1.40 | $0.26 | $4.40 |
| GLM-5.3 Flash | Fast general and agentic workloads | 1,000,000 | $0.15 | $0.03 | $0.50 |
| GLM-5.2 | Reasoning, coding and function calling | 1,000,000 | $1.40 | $0.26 | $4.40 |
| DeepSeek V4 Flash 0731 | General purpose and function calling | 1,000,000 | $0.14 | $0.03 | $0.28 |
| DeepSeek V4 Pro 0813 | Complex reasoning and coding | 1,048,576 | $1.32 | $0.13 | $3.96 |
| MiniMax M3 | General multimodal and agent workflows | 524,288 | $0.30 | $0.06 | $1.20 |
| Inkling | Long-context reasoning and structured output | 524,288 | $1.00 | $0.17 | $4.05 |
| Inkling Small | Efficient long-context work | 524,288 | $0.50 | — | $1.20 |
| GPT-OSS 120B | Open-weight reasoning and tools | 128,000 | $0.15 | — | $0.60 |
| GPT-OSS 20B | Small, fast open-weight workloads | 128,000 | $0.05 | — | $0.20 |
| Qwen3.5 9B | Small, fast chat and vision | 262,144 | $0.17 | — | $0.25 |
| Qwen3.7 Max | Long-context general workloads | — | $1.25 | — | $3.75 |
| Qwen3.7 Plus | General and multilingual work | 1,000,000 | $0.32 | — | $1.28 |
| Qwen3.6 Plus | Long-context chat and analysis | 1,000,000 | $0.50 | — | $3.00 |
| Gemma 4 31B Instruct | Fast general use and vision | 262,144 | $0.39 | — | $0.97 |
| Llama 3.3 70B Instruct Turbo | Established general-purpose deployment | 131,072 | $1.04 | — | $1.04 |
| Muse Glimmer 30B | Creative generation and general text | 131,072 | $0.35 | $0.04 | $1.50 |
| Ternary Bonsai 27B | Experimental efficient inference | 262,144 | Free | — | Free |
Published rates apply at standard settings. MP means generated megapixel. Higher resolution or diffusion steps may increase cost.
| Model | Best considered for | Published rate | Unit |
|---|---|---|---|
| GPT Image 2 | General text-to-image and image editing | $0.053 | image |
| GPT Image 1.5 | General image generation and editing | $0.034 | image |
| Gemini 3 Pro Image | High-quality generation; 1K and 2K reference | $0.134 | image |
| Gemini 3.1 Flash Image | Fast image generation and editing | $0.050 | image |
| Flash Image 2.5 | Rapid multimodal image workflows | $0.039 | image |
| Wan 2.6 Image | General generation | $0.030 | image |
| Qwen Image 2.0 | Text rendering and general generation | $0.035 | MP |
| Qwen Image 2.0 Pro | Higher-quality text and visual production | $0.075 | MP |
| Ideogram 4.0 | Design, typography and brand production | $0.060 | MP |
| FLUX.2 Pro | Professional image generation | $0.030 | MP |
| FLUX.2 Dev | Flexible open development workflows | $0.0154 | MP |
| FLUX.2 Flex | Controllable production workflows | $0.030 | MP |
| FLUX.2 Max | Maximum-quality FLUX generation | $0.070 | MP · 50 default steps |
| FLUX.1 Kontext Pro | Instruction-based image editing | $0.040 | MP · 28 default steps |
| FLUX.1 Kontext Max | High-end instruction-based editing | $0.080 | MP · 28 default steps |
| Imagen 4.0 Fast | Fast general generation | $0.020 | MP |
| Imagen 4.0 Preview | General generation | $0.040 | MP |
| Imagen 4.0 Ultra | High-quality generation | $0.060 | MP |
| Seedream 4.0 | General creative production | $0.030 | MP |
| SD XL | Low-cost open image workflows | $0.0019 | MP |
USD per generated video at the listed base setting where published. Actual price varies by model, duration, resolution and audio.
| Model | Mode | Base setting | Published rate |
|---|---|---|---|
| Seedance 2.5 | Text or image to video | model default | from $0.115 |
| Seedance 2.0 | Text or image to video | model default | from $0.160 |
| Seedance 1.0 Lite | Text or image to video | 720p / 5s | $0.140 |
| Seedance 1.0 Pro | Text or image to video | 1080p / 5s | $0.570 |
| Sora 2 | Text or image to video | 720p / 8s | $0.800 |
| Sora 2 Pro | Text or image to video | 1080p / 8s | $2.400 |
| Veo 3.0 Fast | Text or image to video | 1080p / 8s | $0.800 |
| Veo 3.0 Fast + Audio | Video with generated audio | 1080p / 8s | $1.200 |
| Veo 3.0 | Text or image to video | 720p / 8s | $1.600 |
| Veo 3.0 + Audio | Video with generated audio | 720p / 8s | $3.200 |
| Veo 2.0 | Text or image to video | 720p / 5s | $2.500 |
| FLUX 3 | Text or image to video | model default | from $0.170 |
| Kling 2.1 Standard | Text or image to video | 720p / 5s | $0.180 |
| Kling 2.1 Pro | Text or image to video | 1080p / 5s | $0.320 |
| Kling 2.1 Master | Text or image to video | 1080p / 5s | $0.920 |
| PixVerse v5 | Text or image to video | 1080p / 5s | $0.300 |
| Vidu Q1 | Text or image to video | 1080p / 5s | $0.220 |
| MiniMax Hailuo 02 | Text or image to video | 768p / 10s | $0.490 |
| MiniMax 01 Director | Text or image to video | 720p / 5s | $0.280 |
Speech synthesis is billed by characters. Transcription is billed by audio minute.
| Model | Service | Published rate | Unit |
|---|---|---|---|
| Cartesia Sonic 3 | Text to speech | $65.00 | 1M characters |
| Orpheus 3B | Text to speech | $15.00 | 1M characters |
| Kokoro 82M | Text to speech | $4.00 | 1M characters |
| NVIDIA Nemotron 3.5 ASR Streaming | Speech to text | $0.0015 | audio minute |
| NVIDIA Nemotron 3 ASR Streaming | Speech to text | $0.0015 | audio minute |
| NVIDIA Parakeet TDT 0.6B v3 | Speech to text | $0.0015 | audio minute |
| Whisper Large v3 | Speech to text | $0.0015 | audio minute |
Embedding and moderation prices are per 1M tokens. Execution and storage use their stated units.
| Service | Use | Published rate | Unit |
|---|---|---|---|
| Multilingual E5 Large Instruct | Embeddings for retrieval and semantic search | $0.02 | 1M tokens |
| Llama Guard 4 12B | Content and policy classification | $0.20 | 1M tokens |
| Code Interpreter | Secure execution of model-generated code | $0.03 | 60-minute session |
| Code Sandbox · vCPU | Custom development sandboxes | $0.0446 | vCPU hour |
| Code Sandbox · memory | Memory for custom development sandboxes | $0.0149 | GiB RAM hour |
| Shared filesystem | High-bandwidth storage beside compute | $0.16 | GiB month |
Governing terms
The published rates are effective August 29, 2026. Catalogue, availability and prices may change. An accepted Palace Ring Order or institutional agreement controls.
Taxes, charges outside the listed usage unit, support, implementation, networking, data movement and institution-specific controls may be additional.
Each model remains subject to its applicable license, use restrictions and intellectual-property conditions. Model access grants usage rights under those terms; the provider retains its weights.
Rights and responsibilities for prompts, inputs, outputs and service logs are fixed in the applicable Order, Terms of Service and Data Processing Agreement.