Nemotron 3 Ultra 550b
nemotron-3-ultraNVIDIA's most capable model to date, Nemotron 3 Ultra packs 550 billion total parameters (55 billion active) in an open, efficient hybrid Mamba-Transformer mixture-of-experts design with 1M-token context, excelling at agentic reasoning, coding, planning, and tool use.
Best for
Characteristics
Modalities
Use cases
Anthropic
Yes
Chat Completions
Responses Api
Yes
Provider
NVIDIA
Context Window
1M
Prices shown per 1M tokens.
Input
Rp. 1
$0 USD
Output
Rp. 1
$0 USD
Cache
-
-
All prices are estimates and subject to change.
Similar Models
Related models selected by the catalog administrator.
qwen3.8-flash-freeQwen3.8 Flash Free is a no-cost tier of Qwen3.8 Flash, Alibaba's fast and economical Qwen3.8 multimodal model, giving free access to high-throughput coding, vision, and agentic reasoning workloads with a reduced context budget.
muse-spark-1.2-contributorMuse Spark 1.2 Contributor is Meta's Muse Spark 1.2 reasoning model served through the discounted 'contributor' data-contribution tier, giving users lower pricing in exchange for contributing data-development effort.
qwen3.8-flashQwen3.8 Flash is the official fast-tier API version of Alibaba's Qwen3.8 generation, based on the Qwen3.8-Flash-Next architecture (an experimental Qwen4-preview design) and offering low-latency, high-throughput multimodal reasoning with 1M-token context and built-in tools.