[metadata]
description: DeepInfra startup program. Get access to up to 1B tokens for free inference. Apply now! With DeepInfra, it's easy to deploy the latest state-of-the-art AI models in production
hostname: https://deepinfra.com
keywords: inference,ml,ai,whisper,stable,diffusion
og:description: DeepInfra startup program. Get access to up to 1B tokens for free inference. Apply now! With DeepInfra, it's easy to deploy the latest state-of-the-art AI models in production
og:image: https://deepinfra.com/open-graph.png
og:image:type: image/png
og:locale: en_US
og:site_name: DeepInfra
og:title: DeepStart | Production-Ready Machine Learning Models | DeepInfra
og:url: https://deepinfra.com/deepstart
theme-color: #2A3275
twitter:card: summary_large_image
twitter:description: DeepInfra startup program. Get access to up to 1B tokens for free inference. Apply now! With DeepInfra, it's easy to deploy the latest state-of-the-art AI models in production
twitter:image: https://deepinfra.com/open-graph.png
twitter:site: @DeepInfra
twitter:title: DeepStart | Production-Ready Machine Learning Models | DeepInfra
viewport: initial-scale=1, width=device-width

[canonical-links]
https://deepinfra.com/deepstart

[document-links]
/
/Claude: /claude
/DeepSeek: /deepseek
/Flux: /flux
/Gemini: /gemini
/Llama: /llama
/Mistral: /mistral
/Nemotron: /nemotron
/Qwen: /qwen
About: /about
Automatic Speech Recognition: /models/automatic-speech-recognition
Blog: /blog
Careers: /careers
Chat: /chat
Compare: /compare
Contact Sales: /contact-sales
Contact us: /contact-sales
DeepCluster: /deepcluster
DeepGPT: https://deepgpt.com
DeepStart: /deepstart
Docs: https://docs.deepinfra.com
Embeddings: /models/embeddings
FastVideo / FastWan2.2-TI2V-5B-FullAttn-Diffusers: /FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers
GPUs: /gpu-instances
Log In: /login
Media Center: /media-center
Models: /models
Pricing: /pricing
Privacy Policy: /privacy
Qwen / Qwen3.5-397B-A17B: /Qwen/Qwen3.5-397B-A17B
Qwen / Qwen3.6-35B-A3B: /Qwen/Qwen3.6-35B-A3B
Qwen / Qwen3.8-Max: /Qwen/Qwen3.8-Max
Reranker: /models/reranker
Terms of Service: /terms
Text Generation: /models/text-generation
Text To Image: /models/text-to-image
Text To Music: /models/text-to-music
Text To Speech: /models/text-to-speech
Text To Video: /models/text-to-video
Trust Center: https://trust.deepinfra.com
View all models: /models
View full collection (100+): /models
World Model: /models/world-model
Zero Shot Image Classification: /models/zero-shot-image-classification
https://discord.gg/x88dCvhqYq
https://github.com/DeepInfra
https://linkedin.com/company/deep-infra
https://x.com/DeepInfra
inclusionAI / Ling-3.0-flash: /inclusionAI/Ling-3.0-flash
meta-models / Muse-Glimmer-30B: /meta-models/Muse-Glimmer-30B
moonshotai / Kimi-K2.6: /moonshotai/Kimi-K2.6
nvidia / NVIDIA-Nemotron-3-Super-120B-A12B: /nvidia/NVIDIA-Nemotron-3-Super-120B-A12B
nvidia / NVIDIA-Nemotron-3.5-Lightning: /nvidia/NVIDIA-Nemotron-3.5-Lightning
read the announcement: /series-b
text-generation Qwen/ Qwen3.5-397B-A17B Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks. Priority Flex fp8 256k $0.22 cached, $0.45 in, $3.00 out / 1M: /Qwen/Qwen3.5-397B-A17B
text-generation Qwen/ Qwen3.6-35B-A3B Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. Priority Flex fp8 256k $0.10 in, $0.95 out / 1M: /Qwen/Qwen3.6-35B-A3B
text-generation XiaomiMiMo/ MiMo-V2.5-Pro MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash). Priority Flex fp8 1024k $0.20 cached, $1.00 in, $3.00 out / 1M: /XiaomiMiMo/MiMo-V2.5-Pro
text-generation deepseek-ai/ DeepSeek-V4-Flash DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks. Priority Flex fp4 1024k $0.018 cached, $0.09 in, $0.18 out / 1M: /deepseek-ai/DeepSeek-V4-Flash
text-generation deepseek-ai/ DeepSeek-V4-Flash-0731 DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. Priority Flex fp4 1024k $0.016 cached, $0.08 in, $0.18 out / 1M: /deepseek-ai/DeepSeek-V4-Flash-0731
text-generation deepseek-ai/ DeepSeek-V4-Pro DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks. Priority Flex fp4 1024k $0.10 cached, $1.30 in, $2.60 out / 1M: /deepseek-ai/DeepSeek-V4-Pro
text-generation google/ gemma-4-26B-A4B-it Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output. Priority Flex fp8 256k $0.07 in, $0.34 out / 1M: /google/gemma-4-26B-A4B-it
text-generation moonshotai/ Kimi-K2.6 Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. Priority Flex fp4 256k $0.15 cached, $0.75 in, $3.50 out / 1M: /moonshotai/Kimi-K2.6
text-generation moonshotai/ Kimi-K2.7-Code Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6. Priority Flex Cache retention fp4 256k $0.136 cached, $0.68 in, $3.40 out / 1M: /moonshotai/Kimi-K2.7-Code
text-generation nvidia/ NVIDIA-Nemotron-3-Ultra-550B-A55B Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context. Priority Flex Cache retention fp4 256k $0.10 cached, $0.50 in, $2.20 out / 1M: /nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
text-generation zai-org/ GLM-5.1 GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks). Flex fp4 198k $0.205 cached, $1.05 in, $3.50 out / 1M: /zai-org/GLM-5.1
text-generation zai-org/ GLM-5.2 GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**. Priority Flex fp4 1024k $0.14 cached, $0.75 in, $2.40 out / 1M: /zai-org/GLM-5.2
zai-org / GLM-4.7-Flash: /zai-org/GLM-4.7-Flash

[content]
DeepStart | Production-Ready Machine Learning Models | DeepInfra
We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…
Accept
Reject
DeepInfra raises $107M Series B to scale the inference cloud —
read the announcement
Models
By Category
Automatic Speech Recognition
Embeddings
Reranker
Text Generation
Text To Image
Text To Music
Text To Speech
Text To Video
World Model
Zero Shot Image Classification
View all models
By Family
/Claude
/DeepSeek
/Flux
/Gemini
/Llama
/Mistral
/Nemotron
/Qwen
Docs
Pricing
GPUs
Chat
DeepStart
DeepCluster
Blog
Contact Sales
Log In
Models
Automatic Speech Recognition
Embeddings
Reranker
Text Generation
Text To Image
Text To Music
Text To Speech
Text To Video
World Model
Zero Shot Image Classification
Docs
Pricing
GPUs
Chat
DeepStart
DeepCluster
Blog
Feedback
Contact Sales
Log In
DeepStart
1,000,000,000
free tokens*
To qualify your startup needs to meet the following
Have raised between 250K and 10M USD
Founded in the last 2 years
* at DeepSeek V3.1 prices
Apply now
Featured models
What we loved, used, and implemented most last month
View All
text-generation
text-to-speech
text-to-image
text-generation
deepseek-ai/
DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
Priority
Flex
fp4
1024k
$0.016 cached, $0.08 in, $0.18 out / 1M
text-generation
zai-org/
GLM-5.2
GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.
Priority
Flex
fp4
1024k
$0.14 cached, $0.75 in, $2.40 out / 1M
text-generation
moonshotai/
Kimi-K2.7-Code
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
Priority
Flex
Cache retention
fp4
256k
$0.136 cached, $0.68 in, $3.40 out / 1M
text-generation
nvidia/
NVIDIA-Nemotron-3-Ultra-550B-A55B
Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.
Priority
Flex
Cache retention
fp4
256k
$0.10 cached, $0.50 in, $2.20 out / 1M
text-generation
deepseek-ai/
DeepSeek-V4-Flash
DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.
Priority
Flex
fp4
1024k
$0.018 cached, $0.09 in, $0.18 out / 1M
text-generation
deepseek-ai/
DeepSeek-V4-Pro
DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.
Priority
Flex
fp4
1024k
$0.10 cached, $1.30 in, $2.60 out / 1M
text-generation
moonshotai/
Kimi-K2.6
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
Priority
Flex
fp4
256k
$0.15 cached, $0.75 in, $3.50 out / 1M
text-generation
XiaomiMiMo/
MiMo-V2.5-Pro
MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).
Priority
Flex
fp8
1024k
$0.20 cached, $1.00 in, $3.00 out / 1M
text-generation
Qwen/
Qwen3.6-35B-A3B
Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.
Priority
Flex
fp8
256k
$0.10 in, $0.95 out / 1M
text-generation
zai-org/
GLM-5.1
GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).
Flex
fp4
198k
$0.205 cached, $1.05 in, $3.50 out / 1M
text-generation
Qwen/
Qwen3.5-397B-A17B
Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.
Priority
Flex
fp8
256k
$0.22 cached, $0.45 in, $3.00 out / 1M
text-generation
google/
gemma-4-26B-A4B-it
Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.
Priority
Flex
fp8
256k
$0.07 in, $0.34 out / 1M
View full collection (100+)
Have questions or need a custom solution?
Contact Sales
Company
Pricing
Docs
Compare
DeepStart
About
Careers
Contact us
Media Center
Trust Center
DeepGPT
Latest Models
meta-models
/
Muse-Glimmer-30B
inclusionAI
/
Ling-3.0-flash
FastVideo
/
FastWan2.2-TI2V-5B-FullAttn-Diffusers
nvidia
/
NVIDIA-Nemotron-3.5-Lightning
Qwen
/
Qwen3.8-Max
Featured Models
nvidia
/
NVIDIA-Nemotron-3-Super-120B-A12B
Qwen
/
Qwen3.5-397B-A17B
zai-org
/
GLM-4.7-Flash
Qwen
/
Qwen3.6-35B-A3B
moonshotai
/
Kimi-K2.6
© 2026 DeepInfra. All rights reserved.
Privacy Policy
Terms of Service
