[metadata]
description: DeepInfra offers cost-effective, scalable, easy-to-deploy, and production-ready machine-learning models and infrastructures for deep-learning models.
hostname: https://deepinfra.com
keywords: inference,ml,ai,whisper,stable,diffusion
og:description: DeepInfra offers cost-effective, scalable, easy-to-deploy, and production-ready machine-learning models and infrastructures for deep-learning models.
og:image: https://deepinfra.com/open-graph.png
og:image:type: image/png
og:locale: en_US
og:site_name: DeepInfra
og:title: Machine Learning Models and Infrastructure | DeepInfra
og:url: https://deepinfra.com/
theme-color: #2A3275
twitter:card: summary_large_image
twitter:description: DeepInfra offers cost-effective, scalable, easy-to-deploy, and production-ready machine-learning models and infrastructures for deep-learning models.
twitter:image: https://deepinfra.com/open-graph.png
twitter:site: @DeepInfra
twitter:title: Machine Learning Models and Infrastructure | DeepInfra
viewport: initial-scale=1, width=device-width

[canonical-links]
https://deepinfra.com/

[document-links]
/
/Claude: /claude
/DeepSeek: /deepseek
/Flux: /flux
/Gemini: /gemini
/Llama: /llama
/Mistral: /mistral
/Nemotron: /nemotron
/Qwen: /qwen
About: /about
Announcement · May the 4th, 2026 We've raised $107M Series B to scale the inference cloud Co-led by 500 Global and Georges Harik, with participation from A.Capital Ventures, Crescent Cove, Felicis, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90. Read the announcement May the 4th be with you. $107M Series B funding 25x token volume since Series A: /series-b
Automatic Speech Recognition: /models/automatic-speech-recognition
Blog: /blog
Book a consultation: /contact-sales
Careers: /careers
Chat: /chat
Compare: /compare
Contact Sales: /contact-sales
Contact us: /contact-sales
DeepCluster: /deepcluster
DeepGPT: https://deepgpt.com
DeepStart: /deepstart
Docs: https://docs.deepinfra.com
Embeddings: /models/embeddings
GPUs: /gpu-instances
Learn more: /deepcluster
Let's Go: /dash
Log In: /login
Media Center: /media-center
Models: /models
Pricing: /pricing
Privacy Policy: /privacy
Qwen / Qwen3.8-Max: /Qwen/Qwen3.8-Max
Qwen text-generation Qwen3.5-397B-A17B $0.45/M in • $3.00/M out: /Qwen/Qwen3.5-397B-A17B
Qwen text-generation Qwen3.6-35B-A3B $0.10/M in • $0.95/M out: /Qwen/Qwen3.6-35B-A3B
Reranker: /models/reranker
Terms of Service: /terms
Text Generation: /models/text-generation
Text To Image: /models/text-to-image
Text To Music: /models/text-to-music
Text To Speech: /models/text-to-speech
Text To Video: /models/text-to-video
Trust Center: https://trust.deepinfra.com
View all models: /models
View full collection (100+): /models
World Model: /models/world-model
XiaomiMiMo text-generation MiMo-V2.5-Pro $1.00/M in • $3.00/M out: /XiaomiMiMo/MiMo-V2.5-Pro
Zero Shot Image Classification: /models/zero-shot-image-classification
black-forest-labs / FLUX-2-klein-4b: /black-forest-labs/FLUX-2-klein-4b
deepseek-ai / DeepSeek-V4-Flash-0731: /deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-ai / DeepSeek-V4-Pro: /deepseek-ai/DeepSeek-V4-Pro
deepseek-ai text-generation DeepSeek-V4-Flash $0.09/M in • $0.18/M out: /deepseek-ai/DeepSeek-V4-Flash
deepseek-ai text-generation DeepSeek-V4-Flash-0731 $0.09/M in • $0.18/M out: /deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-ai text-generation DeepSeek-V4-Pro $1.30/M in • $2.60/M out: /deepseek-ai/DeepSeek-V4-Pro
google / gemma-4-26B-A4B-it: /google/gemma-4-26B-A4B-it
google / nano-banana-2-lite: /google/nano-banana-2-lite
google / nano-banana-2: /google/nano-banana-2
google text-generation gemma-4-26B-A4B-it $0.07/M in • $0.34/M out: /google/gemma-4-26B-A4B-it
google text-generation gemma-4-31B-it $0.13/M in • $0.38/M out: /google/gemma-4-31B-it
https://discord.gg/x88dCvhqYq
https://github.com/DeepInfra
https://linkedin.com/company/deep-infra
https://x.com/DeepInfra
moonshotai / Kimi-K2.5: /moonshotai/Kimi-K2.5
moonshotai text-generation Kimi-K2.6 $0.75/M in • $3.50/M out: /moonshotai/Kimi-K2.6
moonshotai text-generation Kimi-K2.7-Code $0.74/M in • $3.50/M out: /moonshotai/Kimi-K2.7-Code
nvidia gpu-rental On-Demand DGX B300 GPUs $4.89 / instance-hour: /gpu-instances
nvidia text-generation NVIDIA-Nemotron-3-Super-120B-A12B $0.085/M in • $0.40/M out: /nvidia/NVIDIA-Nemotron-3-Super-120B-A12B
nvidia text-generation NVIDIA-Nemotron-3-Ultra-550B-A55B $0.50/M in • $2.20/M out: /nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
read the announcement: /series-b
text-generation Qwen/ Qwen3.5-397B-A17B Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks. Priority Flex fp8 256k $0.22 cached, $0.45 in, $3.00 out / 1M: /Qwen/Qwen3.5-397B-A17B
text-generation Qwen/ Qwen3.6-35B-A3B Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. Priority Flex fp8 256k $0.10 in, $0.95 out / 1M: /Qwen/Qwen3.6-35B-A3B
text-generation XiaomiMiMo/ MiMo-V2.5-Pro MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash). Priority Flex fp8 1024k $0.20 cached, $1.00 in, $3.00 out / 1M: /XiaomiMiMo/MiMo-V2.5-Pro
text-generation deepseek-ai/ DeepSeek-V4-Flash DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks. Priority Flex fp4 1024k $0.018 cached, $0.09 in, $0.18 out / 1M: /deepseek-ai/DeepSeek-V4-Flash
text-generation deepseek-ai/ DeepSeek-V4-Flash-0731 DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. Priority Flex fp4 1024k $0.018 cached, $0.09 in, $0.18 out / 1M: /deepseek-ai/DeepSeek-V4-Flash-0731
text-generation deepseek-ai/ DeepSeek-V4-Pro DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks. Priority Flex fp4 1024k $0.10 cached, $1.30 in, $2.60 out / 1M: /deepseek-ai/DeepSeek-V4-Pro
text-generation google/ gemma-4-26B-A4B-it Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output. Priority Flex fp8 256k $0.07 in, $0.34 out / 1M: /google/gemma-4-26B-A4B-it
text-generation moonshotai/ Kimi-K2.6 Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. Priority Flex fp4 256k $0.15 cached, $0.75 in, $3.50 out / 1M: /moonshotai/Kimi-K2.6
text-generation moonshotai/ Kimi-K2.7-Code Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6. Priority Flex fp4 256k $0.15 cached, $0.74 in, $3.50 out / 1M: /moonshotai/Kimi-K2.7-Code
text-generation nvidia/ NVIDIA-Nemotron-3-Ultra-550B-A55B Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context. Priority Flex fp4 256k $0.10 cached, $0.50 in, $2.20 out / 1M: /nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
text-generation zai-org/ GLM-5.1 GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks). Flex fp4 198k $0.205 cached, $1.05 in, $3.50 out / 1M: /zai-org/GLM-5.1
text-generation zai-org/ GLM-5.2 GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**. Priority Flex fp4 1024k $0.14 cached, $0.75 in, $2.40 out / 1M: /zai-org/GLM-5.2
thinkingmachines / Inkling-Small: /thinkingmachines/Inkling-Small
zai-org / GLM-5: /zai-org/GLM-5
zai-org text-generation GLM-5 $0.60/M in • $2.08/M out: /zai-org/GLM-5
zai-org text-generation GLM-5.1 $1.05/M in • $3.50/M out: /zai-org/GLM-5.1
zai-org text-generation GLM-5.2 $0.75/M in • $2.40/M out: /zai-org/GLM-5.2

[content]
Machine Learning Models and Infrastructure | DeepInfra
We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…
Accept
Reject
DeepInfra raises $107M Series B to scale the inference cloud —
read the announcement
Models
By Category
Automatic Speech Recognition
Embeddings
Reranker
Text Generation
Text To Image
Text To Music
Text To Speech
Text To Video
World Model
Zero Shot Image Classification
View all models
By Family
/Claude
/DeepSeek
/Flux
/Gemini
/Llama
/Mistral
/Nemotron
/Qwen
Docs
Pricing
GPUs
Chat
DeepStart
DeepCluster
Blog
Contact Sales
Log In
Models
Automatic Speech Recognition
Embeddings
Reranker
Text Generation
Text To Image
Text To Music
Text To Speech
Text To Video
World Model
Zero Shot Image Classification
Docs
Pricing
GPUs
Chat
DeepStart
DeepCluster
Blog
Feedback
Contact Sales
Log In
FAST
SIMPLE
RELIABLE
LOW-COST
AI Inference
Accelerate your AI with developer-friendly APIs designed for performance and cost-efficiency.
Let's Go
Book a consultation
deepseek-ai
text-generation
DeepSeek-V4-Flash-0731
$0.09/M in • $0.18/M out
zai-org
text-generation
GLM-5.2
$0.75/M in • $2.40/M out
moonshotai
text-generation
Kimi-K2.7-Code
$0.74/M in • $3.50/M out
nvidia
gpu-rental
On-Demand DGX B300 GPUs
$4.89 / instance-hour
nvidia
text-generation
NVIDIA-Nemotron-3-Ultra-550B-A55B
$0.50/M in • $2.20/M out
deepseek-ai
text-generation
DeepSeek-V4-Flash
$0.09/M in • $0.18/M out
deepseek-ai
text-generation
DeepSeek-V4-Pro
$1.30/M in • $2.60/M out
moonshotai
text-generation
Kimi-K2.6
$0.75/M in • $3.50/M out
XiaomiMiMo
text-generation
MiMo-V2.5-Pro
$1.00/M in • $3.00/M out
Qwen
text-generation
Qwen3.6-35B-A3B
$0.10/M in • $0.95/M out
zai-org
text-generation
GLM-5.1
$1.05/M in • $3.50/M out
Qwen
text-generation
Qwen3.5-397B-A17B
$0.45/M in • $3.00/M out
google
text-generation
gemma-4-26B-A4B-it
$0.07/M in • $0.34/M out
google
text-generation
gemma-4-31B-it
$0.13/M in • $0.38/M out
nvidia
text-generation
NVIDIA-Nemotron-3-Super-120B-A12B
$0.085/M in • $0.40/M out
zai-org
text-generation
GLM-5
$0.60/M in • $2.08/M out
deepseek-ai
text-generation
DeepSeek-V4-Flash-0731
$0.09/M in • $0.18/M out
zai-org
text-generation
GLM-5.2
$0.75/M in • $2.40/M out
moonshotai
text-generation
Kimi-K2.7-Code
$0.74/M in • $3.50/M out
nvidia
gpu-rental
On-Demand DGX B300 GPUs
$4.89 / instance-hour
nvidia
text-generation
NVIDIA-Nemotron-3-Ultra-550B-A55B
$0.50/M in • $2.20/M out
deepseek-ai
text-generation
DeepSeek-V4-Flash
$0.09/M in • $0.18/M out
deepseek-ai
text-generation
DeepSeek-V4-Pro
$1.30/M in • $2.60/M out
moonshotai
text-generation
Kimi-K2.6
$0.75/M in • $3.50/M out
XiaomiMiMo
text-generation
MiMo-V2.5-Pro
$1.00/M in • $3.00/M out
Qwen
text-generation
Qwen3.6-35B-A3B
$0.10/M in • $0.95/M out
zai-org
text-generation
GLM-5.1
$1.05/M in • $3.50/M out
Qwen
text-generation
Qwen3.5-397B-A17B
$0.45/M in • $3.00/M out
google
text-generation
gemma-4-26B-A4B-it
$0.07/M in • $0.34/M out
google
text-generation
gemma-4-31B-it
$0.13/M in • $0.38/M out
nvidia
text-generation
NVIDIA-Nemotron-3-Super-120B-A12B
$0.085/M in • $0.40/M out
zai-org
text-generation
GLM-5
$0.60/M in • $2.08/M out
Let's Go
Book a consultation
Announcement · May the 4th, 2026
We've raised $107M Series B
to scale the inference cloud
Co-led by 500 Global and Georges Harik, with participation from A.Capital Ventures, Crescent Cove, Felicis, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90.
Read the announcement
May the 4th be with you.
$107M
Series B funding
25x
token volume since Series A
Scale to trillions of tokens without breaking the bank
Low pay-as-you-go pricing - no long-term contracts, no hidden fees, no surprises. Startup? Enterprise? We can scale. We are there for you with our simple APIs and hands-on technical support.
Inference Tailored to You
An inference partner that meets your needs. Whether you're optimizing for cost, latency, throughput or scale - we design the solution around your priorities. DeepInfra provides 100+ models to cover all your needs.
Zero Retention. Compliant. Secure.
With our zero retention policy your inputs, your outputs, and your user data stay private. DeepInfra is SOC 2 and ISO 27001 certified. We follow the best practices in information security and privacy.
Our Hardware. Our Data Centers. Your Performance Edge.
DeepInfra runs on our own cutting-edge inference optimised infrastructure, in secure US-based data centers. Better performance and reliability for you.
Models
Explore our Featured Models
View All
text-generation
text-to-speech
text-to-image
text-generation
deepseek-ai/
DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
Priority
Flex
fp4
1024k
$0.018 cached, $0.09 in, $0.18 out / 1M
text-generation
zai-org/
GLM-5.2
GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.
Priority
Flex
fp4
1024k
$0.14 cached, $0.75 in, $2.40 out / 1M
text-generation
moonshotai/
Kimi-K2.7-Code
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
Priority
Flex
fp4
256k
$0.15 cached, $0.74 in, $3.50 out / 1M
text-generation
nvidia/
NVIDIA-Nemotron-3-Ultra-550B-A55B
Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.
Priority
Flex
fp4
256k
$0.10 cached, $0.50 in, $2.20 out / 1M
text-generation
deepseek-ai/
DeepSeek-V4-Flash
DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.
Priority
Flex
fp4
1024k
$0.018 cached, $0.09 in, $0.18 out / 1M
text-generation
deepseek-ai/
DeepSeek-V4-Pro
DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.
Priority
Flex
fp4
1024k
$0.10 cached, $1.30 in, $2.60 out / 1M
text-generation
moonshotai/
Kimi-K2.6
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
Priority
Flex
fp4
256k
$0.15 cached, $0.75 in, $3.50 out / 1M
text-generation
XiaomiMiMo/
MiMo-V2.5-Pro
MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).
Priority
Flex
fp8
1024k
$0.20 cached, $1.00 in, $3.00 out / 1M
text-generation
Qwen/
Qwen3.6-35B-A3B
Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.
Priority
Flex
fp8
256k
$0.10 in, $0.95 out / 1M
text-generation
zai-org/
GLM-5.1
GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).
Flex
fp4
198k
$0.205 cached, $1.05 in, $3.50 out / 1M
text-generation
Qwen/
Qwen3.5-397B-A17B
Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.
Priority
Flex
fp8
256k
$0.22 cached, $0.45 in, $3.00 out / 1M
text-generation
google/
gemma-4-26B-A4B-it
Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.
Priority
Flex
fp8
256k
$0.07 in, $0.34 out / 1M
View full collection (100+)
Live AI Inference Metrics
End-to-end insights into speed, scale, stability and spend
0.00
M
Tokens per second
0
ms
Time to first token
0
Requests per second
0.00
exaFLOPS
DeepCluster
Your own NVIDIA B300
GPU cluster
Dedicated hardware, procured and operated by DeepInfra. Full ownership, Tier 3 datacenter, 99.982% uptime SLA.
Learn more
NVIDIA B300 · 5-year term
$1.98
/GPU-hr
vs
$6.50
/GPU-hr on public cloud
70%
cheaper than cloud
288 GB
HBM3e per GPU
256–5,000
GPUs available
Have questions or need a custom solution?
Contact Sales
Company
Pricing
Docs
Compare
DeepStart
About
Careers
Contact us
Media Center
Trust Center
DeepGPT
Latest Models
Qwen
/
Qwen3.8-Max
deepseek-ai
/
DeepSeek-V4-Flash-0731
thinkingmachines
/
Inkling-Small
google
/
nano-banana-2-lite
google
/
nano-banana-2
Featured Models
deepseek-ai
/
DeepSeek-V4-Pro
moonshotai
/
Kimi-K2.5
zai-org
/
GLM-5
google
/
gemma-4-26B-A4B-it
black-forest-labs
/
FLUX-2-klein-4b
© 2026 DeepInfra. All rights reserved.
Privacy Policy
Terms of Service
