Back to AI Models

Nemotron 3 Ultra 550b

nemotron-3-ultra
Try with NaraRouter

NVIDIA's most capable model to date, Nemotron 3 Ultra packs 550 billion total parameters (55 billion active) in an open, efficient hybrid Mamba-Transformer mixture-of-experts design with 1M-token context, excelling at agentic reasoning, coding, planning, and tool use.

Text
Overview

Best for

General assistanceWorkflows

Characteristics

StreamingFast

Modalities

Text

Use cases

General AssistantWorkflow Automation
Endpoint Compatibility

Anthropic

Yes

Chat Completions

Recommended

Responses Api

Yes

Specification

Provider

NVIDIA

Context Window

1M

Pricing

Prices shown per 1M tokens.

Input

Rp. 1

$0 USD

Output

Rp. 1

$0 USD

Cache

-

-

All prices are estimates and subject to change.

Capabilities
ReasoningCodingTool CallingVisionStructured OutputStreaming

Similar Models

Related models selected by the catalog administrator.

Qwen 3.8 Flash Free
qwen3.8-flash-free

Qwen3.8 Flash Free is a no-cost tier of Qwen3.8 Flash, Alibaba's fast and economical Qwen3.8 multimodal model, giving free access to high-throughput coding, vision, and agentic reasoning workloads with a reduced context budget.

TextVision
Muse Spark 1.2 Contributor
muse-spark-1.2-contributor

Muse Spark 1.2 Contributor is Meta's Muse Spark 1.2 reasoning model served through the discounted 'contributor' data-contribution tier, giving users lower pricing in exchange for contributing data-development effort.

TextVision
Qwen 3.8 Flash
qwen3.8-flash

Qwen3.8 Flash is the official fast-tier API version of Alibaba's Qwen3.8 generation, based on the Qwen3.8-Flash-Next architecture (an experimental Qwen4-preview design) and offering low-latency, high-throughput multimodal reasoning with 1M-token context and built-in tools.

TextVision
Nemotron 3 Ultra 550b | NaraRouter