[metadata]
description: Access 200+ AI models with one API. Launch secure agent sandboxes and GPU instances in minutes. Built for developers, priced for startups.
keywords: 4090,gpu,ai,A100,h100
og:description: Novita AI provides 200+ Model APIs, custom deployment, GPU Instances, and Serverless GPUs. Scale AI, optimize performance, and innovate with ease and efficiency.
og:image: https://novita.ai/mainpage/media-logo.png
og:image:alt: Novita AI Logo
og:site_name: Novita AI
og:title: Novita AI – Model Libraries & GPU Cloud - Deploy, Scale & Innovate
og:type: website
og:url: https://novita.ai
twitter:card: summary_large_image
twitter:description: Novita AI provides 200+ Model APIs, custom deployment, GPU Instances, and Serverless GPUs. Scale AI, optimize performance, and innovate with ease and efficiency.
twitter:image: https://novita.ai/mainpage/media-logo.png
twitter:image:alt: Novita AI Logo
twitter:title: Novita AI – Model Libraries & GPU Cloud - Deploy, Scale & Innovate
viewport: width=device-width, initial-scale=1

[canonical-links]
https://novita.ai

[document-links]
/
ANNOUNCEMENT Deepseek V4 Pro Available on Novita AI now LLM $1.6/3.38 in/out MTokens | 1048576 Context: https://novita.ai/models/model-detail/deepseek-deepseek-v4-pro
ANNOUNCEMENT Featured Blogs insights, LLM tips, and GPU solutions CHECK OUT THE LATEST ARTICLES: https://blogs.novita.ai/
ANNOUNCEMENT Ling-3.0-flash Available on Novita AI now LLM $0/0 in/out M Tokens | 262144 Context: https://novita.ai/models/model-detail/inclusionai-ling-3.0-flash
ANNOUNCEMENT Macaron V1 Venti Available on Novita AI now LLM $0/0 in/out MTokens | 1048576 Context: https://novita.ai/models/model-detail/mindai-macaron-v1-venti
Agent Sandbox: /sandbox
Blog: https://blogs.novita.ai
CASE STUDY Novita X Accelerate AI Inference Accelerate AI Inference: https://blogs.novita.ai/sglang/
CASE STUDY Novita X Hugging Face Novita available on Hugging Face now Novita available on Hugging Face now: https://blogs.novita.ai/novita-ai-is-now-available-on-hugging-face/
CASE STUDY Novita X LLM Accelerate AI Inference Accelerate AI Inference: https://blogs.novita.ai/announcing-our-partnership-with-vllm-to-advance-ai-inference
CASE STUDY Novita X POE Novita models on POE now Novita models on POE now: https://poe.com/novitaai
Careers: https://jobs.ashbyhq.com/novita-ai
Contact Support: mailto:support@novita.ai
Cookie Policy: /legal/cookie-policy
DeepSeek V4 Pro $1.74/Mt Input · $3.48/Mt Output 1M Context LLM: /models/model-detail/deepseek-deepseek-v4-pro
Docs: https://docs.novita.ai/guides/introduction
Explore All Models: /models
GLM 5.1 $1.4/Mt Input · $4.4/Mt Output 200K Context LLM: /models/model-detail/zai-org-glm-5.1
GPU Bare Metal: /gpu-baremetal
GPU Instance: /gpus
Gemma 4 31B $0.14/Mt Input · $0.4/Mt Output 256K Context LLM: /models/model-detail/google-gemma-4-31b-it
Get Started: /dedicated-endpoint
Get Started: /sandbox
Kimi K2.6 $0.95/Mt Input · $4/Mt Output 256K Context LLM: /models/model-detail/moonshotai-kimi-k2.6
MiniMax M2.7 $0.3/Mt Input · $1.2/Mt Output 200K Context LLM: /models/model-detail/minimax-minimax-m2.7
Model APIs: /models
Pricing: /pricing
Privacy Policy: /legal/privacy-policy
Qwen3.5 397B A17B $0.6/Mt Input · $3.6/Mt Output 256K Context LLM: /models/model-detail/qwen-qwen3.5-397b-a17b
Refer to Earn: /affiliate-new
Start Building: /user/register
Supply GPUs: mailto:gpu@novita.ai
Talk to Sales: https://meetings-na2.hubspot.com/junyu
Talk to Us: https://meetings-na2.hubspot.com/junyu
Terms of Service: /legal/terms-of-service
Trust Center: https://trust.novita.ai
https://docs.novita.ai/skill.md: https://docs.novita.ai/skill.md
https://www.linkedin.com/company/novita-ai-labs/
https://x.com/novita_labs
llms.txt: /llms.txt

[content]
Novita AI - AI & Agent Cloud for Developers
For agents: fetch the complete documentation index at
llms.txt
. Markdown is available with
Accept: text/markdown
and with
.md
URL variants.
Model APIs
Agent Sandbox
GPUs
Resources
Pricing
Start Building
Model APIs
Agent Sandbox
GPU Cloud
The AI-Native Cloud
for Builders and
Agents
Run models, scale GPUs, and build AI agents, all on one platform.
Start Building
For Agent
Click to copy instructions for you agent:
Read
https://docs.novita.ai/skill.md
and follow the instructions.
Talk to Us
Trusted by
Model APIs
Agent Sandbox
GPU Cloud
MODEL APIS
LLM
IMAGE
AUDIO
VIDEO
VISION
MODEL
"KIMI-K2.5"
200+
models
200ms
latency
99.5%
uptime
1
Serverless Model APIs
Run 200+ models through a single API.
No infrastructure to manage.
Text, image, audio, video — all serverless, all
production-ready. You call it, we run it. Billed by the
token, not the hour.
Explore All Models
DeepSeek V4 Pro
$1.74/Mt Input · $3.48/Mt Output
1M Context
LLM
MiniMax M2.7
$0.3/Mt Input · $1.2/Mt Output
200K Context
LLM
GLM 5.1
$1.4/Mt Input · $4.4/Mt Output
200K Context
LLM
Kimi K2.6
$0.95/Mt Input · $4/Mt Output
256K Context
LLM
Gemma 4 31B
$0.14/Mt Input · $0.4/Mt Output
256K Context
LLM
Qwen3.5 397B A17B
$0.6/Mt Input · $3.6/Mt Output
256K Context
LLM
DeepSeek V4 Pro
$1.74/Mt Input · $3.48/Mt Output
1M Context
LLM
MiniMax M2.7
$0.3/Mt Input · $1.2/Mt Output
200K Context
LLM
GLM 5.1
$1.4/Mt Input · $4.4/Mt Output
200K Context
LLM
Kimi K2.6
$0.95/Mt Input · $4/Mt Output
256K Context
LLM
Gemma 4 31B
$0.14/Mt Input · $0.4/Mt Output
256K Context
LLM
Qwen3.5 397B A17B
$0.6/Mt Input · $3.6/Mt Output
256K Context
LLM
DeepSeek V4 Pro
$1.74/Mt Input · $3.48/Mt Output
1M Context
LLM
MiniMax M2.7
$0.3/Mt Input · $1.2/Mt Output
200K Context
LLM
GLM 5.1
$1.4/Mt Input · $4.4/Mt Output
200K Context
LLM
Kimi K2.6
$0.95/Mt Input · $4/Mt Output
256K Context
LLM
Gemma 4 31B
$0.14/Mt Input · $0.4/Mt Output
256K Context
LLM
Qwen3.5 397B A17B
$0.6/Mt Input · $3.6/Mt Output
256K Context
LLM
2
Dedicated Endpoints
Private endpoints. Guaranteed performance. No noisy neighbors.
Your model. Your compute. Isolated resources mean consistent latency at any throughput. Because production doesn't have a retry budget.
Get Started
BASE_URL
"API.NOVITA.AI/YOUR-ENDPOINT"
OPERATIONAL
AGENT SANDBOX
agent
"coding agents"
"coding agents"
coding agent · active
sandbox runtime
Run test suite · pytest
queued
Write fix · patch applied
running
Identify bug · null pointer line 84
done
Read codebase · src/api/routes.py
done
startup
~200ms
isolation
Full
billing
per second
status
RUNNING
1
Agent sandbox
Secure, isolated runtimes. Built for agents that actually do things.
Not a notebook. Not a container you configure yourself. A purpose-built environment where agents run, use tools, call models, and execute tasks — cleanly, in isolation, every time.
Get Started
GPU CLOUD
GPU
flagship
flagship
1
GPU Instances
Full-control GPU machines. Yours in seconds.
Deploy models, run inference, train from scratch, on dedicated GPU instances you fully control. Predictable performance. No shared resources. No surprises.
Get Started
2
Serverless GPU
Submit a job. We handle the rest.
No instances to provision. No idle compute to pay for. Novita allocates GPU resources automatically, scales up under load, scales to zero when you're done. You pay for execution, nothing else.
Get Started
job
queued
running
complete
allocating gpu resources
processing job
processing job
allocating
running
complete
12
%
allocated
auto
duration
0.1s
cost
$0.0001
idle time
$0.00
cluster
"Cluster-01"
"Cluster-01"
CLUSTER-01 · 6 nodes
NVLink · GPUDirect RDMA · PCIe
CLUSTER-02 · 12 nodes
NVLink · GPUDirect RDMA · PCIe
CLUSTER-03 · 3 nodes
PCIe · 10GbE
Node-01
51%
Node-02
79%
Node-03
86%
Node-05
89%
Node-06
65%
Node-07
81%
GPU
8× NVIDIA H200
GPU Memory
141 GB HBM3e per GPU
1.128
TB total
Nodes
6 / 6
Interconnect
NVLink 4th Gen · 900 GB/s
Network
400 Gb/s RDMA
GPU
8× NVIDIA H100
GPU Memory
80 GB HBM3 per GPU
640
GB total
Nodes
6/6
Interconnect
NVLink 4th Gen · 900 GB/s
Network
400 Gb/s RDMA
GPU
8× NVIDIA H100
GPU Memory
80 GB HBM3 per GPU
640
GB total
Nodes
6/6
Interconnect
NVLink 4th Gen · 900 GB/s
Network
400 Gb/s RDMA
3
Bare Metal
Maximum performance. Zero abstraction overhead.
Dedicated physical GPU clusters for large-scale inference, training runs, and enterprise deployments that can't compromise on throughput. When you need the hardware to yourself, this is it.
Get Started
Why Novita AI
Built for AI from day one. Designed for what you're actually building.
Get Started
Better price-performance
Up to 50% less than major cloud providers. Not because we cut corners, because we built the infrastructure.
Built for production reliability
Stable infrastructure with low latency, high throughput, and reliable uptime at scale.
One platform for the full AI stack
Model APIs, GPU infrastructure, and agent runtimes — all in one platform.
Scale with your workload
Start small and scale seamlessly from APIs to dedicated clusters.
Dedicated support when it matters
Fast technical support from a team that understands AI infrastructure.
Built with Novita AI
Testimonials
Don't take our word for it.
"
I appreciate how fast Novita AI moves to deploy newly released models. Their team is often the first to get stable, production ready inference support online – often on Day One. That speed is critical for the whole open-source AI community.
"
Julien Chaumond
Co-Founder & CTO
"
Novita has been a huge help for us at Fish Audio. Their reliable GPU infrastructure allows us focus on developing and improving our text-to-speech models instead of dealing with hardware headaches. Their support and performance have made it much easier to push our work forward.
"
Shijia Liao
Co-Founder & Chief Scientist
"
Novita's Model API was super simple to integrate, and it's been great in powering our AI-driven flashcards and quizzes. The platform takes care of the heavy lifting, so we can focus on building better learning tools for our users without worrying about infrastructure or scaling issues.
"
Petros Christodoulou
Co-Founder and CEO
"
Working with Novita AI has been a fantastic experience for Kilo. Their inference platform helps us deliver fast and reliable AI coding workflows across multiple LLMs, with strong real-world performance for agentic workflows. And the team has been remarkably easy to work with! They are always optimizing based on the latest models and technology—a perfect partner for Kilo Code.
"
Ari Messer
Head of Partnerships
What's New
ANNOUNCEMENT
Macaron V1 Venti
Available on Novita AI now
LLM
$0/0 in/out MTokens | 1048576 Context
ANNOUNCEMENT
Ling-3.0-flash
Available on Novita AI now
LLM
$0/0 in/out M Tokens | 262144 Context
ANNOUNCEMENT
Deepseek V4 Pro
Available on Novita AI now
LLM
$1.6/3.38 in/out MTokens | 1048576 Context
CASE STUDY
Novita
X
Hugging Face
Novita available on Hugging Face now
Novita available on Hugging Face now
CASE STUDY
Novita
X
POE
Novita models on POE now
Novita models on POE now
CASE STUDY
Novita
X
LLM
Accelerate AI Inference
Accelerate AI Inference
CASE STUDY
Novita
X
Accelerate AI Inference
Accelerate AI Inference
ANNOUNCEMENT
Featured Blogs
insights, LLM tips, and GPU solutions
CHECK OUT THE LATEST ARTICLES
Everything you need to build production AI.
200+ models, on-demand GPUs, and secure agent runtimes — unified under one API. Free to start, scales as you grow.
Get Started
Power your AI applications with Novita AI's model APIs, GPU instances, and agent sandbox.
Product
Model APIs
Agent Sandbox
GPU Instance
GPU Bare Metal
Resource
Talk to Sales
Contact Support
Pricing
Docs
Partners
Refer to Earn
Supply GPUs
Company
Careers
Blog
Trust Center
©
2026
Novita AI. All rights reserved
Terms of Service
Privacy Policy
Cookie Policy
Cookie Settings
Do Not Sell or Share My Personal Information
