{"id":79654,"topic":"ai","source":"NVIDIA Developer","title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","url_hash":"914fe4ca71f275f4c13487695952b9b6bcfd9304","author":"","summary":"<a href=\"https://news.google.com/rss/articles/CBMisAFBVV95cUxQVnVRQktsTzVzYWxfeHV0MXpZaEVZWGc5aURfNVRleFI4S0daT1pDcDV0bG5rVFdNYXJsanF0NEVRQlFFU1E2dE9oQkdfWEk1Tm1aZ0NKU2VjalZMeUVCbXRTb2lnRG5tVG9mZm9xaUxSZ1Z0eTVHYkRTXzhVYVZ2NEFaV0tWV0NzUnl0ZzZmdU9jRmRleElhaHpZRnRlWkhuaXp4YU5TNG1HZjI0bE8tSg?oc=5\" target=\"_blank\">How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">NVIDIA Developer</font>","content":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.\nThat tradeoff matters even more for agentic AI workloads, where prompts can be long, context can be reused across steps, and applications often stream extended responses back to users.\nNVIDIA NIM packages model- and GPU-aware serving choices into a deployable microservice. Instead of starting from a blank runtime configuration, developers get a validated serving configuration and a supported deployment path, while retaining the ability to benchmark the NIM against their own traffic.\nInference performance is a system property. Precision and kernels, parallelism, scheduling, batching, memory allocation, prefix reuse, model-specific state caches, and decoding strategy all interact. A configuration is useful only if it improves throughput while staying within the application latency target.\nNIM turns this optimization work into a tested starting point. NVIDIA engineers validate configurations for supported model, GPU, and precision combinations, then package the runtime and model artifacts behind standard APIs. For production deployments, NIM Certified adds regular inference-stack updates, CVE handling, broader hardware validation, and commercial support through NVIDIA AI Enterprise.\nThe benchmark definition used throughout this article is:\n- Hardware: 4xB200\n- Agentic workload: 64K/400/76% KV reuse/50 TPS/user (20 ms ITL)\nThis Pareto chart compares the open-source baseline serving stack (NIM Off) against the fully optimized NIM 2.0.12 serving stack (NIM On).\n| Configuration | Native 256K max context | What it represents | \n|---|---|---|\n| NIM Off baseline | 718 tok/s | No NIM optimizations | \n| NIM On (2.0.12) optimized serving stack | 1,997 tok/s 2.5x v. Baseline | Optimized NIM stack: cache/state reuse, MTP speculative decoding and associated fixes, autotuned kernels, partial-prefix matching, scheduler, batching, memory, and parallelism tuning | \nThe measured gains come from interacting configuration bundles, not independent switches whose percentages can simply be added. The main optimization layers are:\n- Precision and autotuned model-aware kernels. Autotuned mixture-of-experts and Mamba kernels map the hybrid architecture efficiently to NVIDIA Blackwell GPUs.\n- Parallel execution. Tensor parallelism distributes the model across four GPUs, while expert-aware execution improves utilization for the mixture-of-experts layers.\n- Prefix and model-state reuse. Prefix caching avoids recomputing repeated context, partial-prefix matching recovers reuse when only part of a prefix matches, and Mamba state-cache settings are tuned for the model architecture.\n- Scheduler, batching, and memory tuning. Concurrent-sequence limits, batched-token limits, block size, and GPU-memory allocation keep more work in flight without crossing the latency target.\n- MTP speculative decoding. NIM 2.0.12 optimized serving stack (including MTP) adds MTP and its associated fixes to the same optimized serving stack. The incremental benefit depends on acceptance rate and available memory headroom.\nThe published curves are a starting point, not a promise that every application will see the same result. The fastest way to determine fit is to replay representative traffic and build a Pareto curve for the latency metric that matters to your users.\n- Deploy the exact software versions. Use NIM 2.0.12 (or newer version), and pin the image tag or digest for every run.\n- Prepare representative traffic. Use a Mooncake-format JSONL trace or capture controlled NIM requests, with appropriate access controls and sanitization for sensitive data.\n- Measure performance. Use NVIDIA AIPerf to replay representative traffic and get perf benchmarks\n- Select the Pareto point that meets the SLO. Compare output throughput among points that satisfy the SLO/latency constraints and determine fit for deployment:\nfor C in 1 4 8 16 32 64; do\n   aiperf profile \\\n \t--model nvidia/nemotron-3-ultra-550b-a55b \\\n \t--endpoint-type chat --streaming \\\n \t--url localhost:8000 \\\n \t--input-file ./agentic-trace.jsonl \\\n \t--custom-dataset-type mooncake_trace \\\n \t--no-fixed-schedule \\\n \t--concurrency \"$C\"\n done\nExample AIPerf concurrency sweep. Replace the trace, request count, and endpoint details with the workload you want to model.\nStart from the Nemotron 3 Ultra NIM page, accept the governing terms, and select the NIM 2.0.12 tag or the exact published digest. After downloading, find and select a profile. The following lists all profiles packaged in the NIM:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG list-model-profiles\nTo pick an optimized profile for agentic workloads on a four-GPU B200 system, you would set the NIM_MODEL_PROFILE to vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0 and enable speculative decoding:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n export NIM_MODEL_PROFILE=vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -e NIM_SPECDEC_ENABLE=1 \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG\nNVIDIA NIM packages model- and GPU-aware serving optimizations into a deployable microservice. For Nemotron 3 Ultra on a 4xB200 system, these optimizations let organizations serve up to 2.5x more users at 50 TPS/user compared with a NIM-off baseline, while keeping the deployment path practical for real multi-user serving.\nNIM combines validated performance engineering with regular inference-stack updates, CVE handling, hardware validation, and commercial support through NVIDIA AI Enterprise. More performance-optimized NIM configurations are planned across a broader range of models, giving developers additional validated starting points for their own latency, throughput, and cost objectives.\nDownload the Nemotron 3 Ultra NIM 2.0.12 from NGC, run it on your NVIDIA GPU infrastructure, and replay representative traffic with NVIDIA AIPerf to select the Pareto point that meets your application target.\nResources","image_url":"https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image4_1480x833-660x370.jpg","lang":"en","published_at":"2026-09-10T17:00:06+00:00","fetched_at":"2026-09-10T18:15:05+00:00","status":"read","starred":0,"extract_state":"ok","summary_auto":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.","cluster_id":null,"extract_retries":0,"extract_error":null,"contract_version":"news_item.v1","format_contract_version":"news_item_formats.v1","dedup_url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 6835 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":6835,"summary_length":264,"usable_text_length":6835,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":6835,"summary_length":264}},"news_item":{"id":79654,"canonical_url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","source_url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","source_name":"NVIDIA Developer","author":null,"published_at":"2026-09-10T17:00:06+00:00","locale":"en","topic":"ai","tags":[],"rss_summary":"<a href=\"https://news.google.com/rss/articles/CBMisAFBVV95cUxQVnVRQktsTzVzYWxfeHV0MXpZaEVZWGc5aURfNVRleFI4S0daT1pDcDV0bG5rVFdNYXJsanF0NEVRQlFFU1E2dE9oQkdfWEk1Tm1aZ0NKU2VjalZMeUVCbXRTb2lnRG5tVG9mZm9xaUxSZ1Z0eTVHYkRTXzhVYVZ2NEFaV0tWV0NzUnl0ZzZmdU9jRmRleElhaHpZRnRlWkhuaXp4YU5TNG1HZjI0bE8tSg?oc=5\" target=\"_blank\">How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">NVIDIA Developer</font>","full_text":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.\nThat tradeoff matters even more for agentic AI workloads, where prompts can be long, context can be reused across steps, and applications often stream extended responses back to users.\nNVIDIA NIM packages model- and GPU-aware serving choices into a deployable microservice. Instead of starting from a blank runtime configuration, developers get a validated serving configuration and a supported deployment path, while retaining the ability to benchmark the NIM against their own traffic.\nInference performance is a system property. Precision and kernels, parallelism, scheduling, batching, memory allocation, prefix reuse, model-specific state caches, and decoding strategy all interact. A configuration is useful only if it improves throughput while staying within the application latency target.\nNIM turns this optimization work into a tested starting point. NVIDIA engineers validate configurations for supported model, GPU, and precision combinations, then package the runtime and model artifacts behind standard APIs. For production deployments, NIM Certified adds regular inference-stack updates, CVE handling, broader hardware validation, and commercial support through NVIDIA AI Enterprise.\nThe benchmark definition used throughout this article is:\n- Hardware: 4xB200\n- Agentic workload: 64K/400/76% KV reuse/50 TPS/user (20 ms ITL)\nThis Pareto chart compares the open-source baseline serving stack (NIM Off) against the fully optimized NIM 2.0.12 serving stack (NIM On).\n| Configuration | Native 256K max context | What it represents | \n|---|---|---|\n| NIM Off baseline | 718 tok/s | No NIM optimizations | \n| NIM On (2.0.12) optimized serving stack | 1,997 tok/s 2.5x v. Baseline | Optimized NIM stack: cache/state reuse, MTP speculative decoding and associated fixes, autotuned kernels, partial-prefix matching, scheduler, batching, memory, and parallelism tuning | \nThe measured gains come from interacting configuration bundles, not independent switches whose percentages can simply be added. The main optimization layers are:\n- Precision and autotuned model-aware kernels. Autotuned mixture-of-experts and Mamba kernels map the hybrid architecture efficiently to NVIDIA Blackwell GPUs.\n- Parallel execution. Tensor parallelism distributes the model across four GPUs, while expert-aware execution improves utilization for the mixture-of-experts layers.\n- Prefix and model-state reuse. Prefix caching avoids recomputing repeated context, partial-prefix matching recovers reuse when only part of a prefix matches, and Mamba state-cache settings are tuned for the model architecture.\n- Scheduler, batching, and memory tuning. Concurrent-sequence limits, batched-token limits, block size, and GPU-memory allocation keep more work in flight without crossing the latency target.\n- MTP speculative decoding. NIM 2.0.12 optimized serving stack (including MTP) adds MTP and its associated fixes to the same optimized serving stack. The incremental benefit depends on acceptance rate and available memory headroom.\nThe published curves are a starting point, not a promise that every application will see the same result. The fastest way to determine fit is to replay representative traffic and build a Pareto curve for the latency metric that matters to your users.\n- Deploy the exact software versions. Use NIM 2.0.12 (or newer version), and pin the image tag or digest for every run.\n- Prepare representative traffic. Use a Mooncake-format JSONL trace or capture controlled NIM requests, with appropriate access controls and sanitization for sensitive data.\n- Measure performance. Use NVIDIA AIPerf to replay representative traffic and get perf benchmarks\n- Select the Pareto point that meets the SLO. Compare output throughput among points that satisfy the SLO/latency constraints and determine fit for deployment:\nfor C in 1 4 8 16 32 64; do\n   aiperf profile \\\n \t--model nvidia/nemotron-3-ultra-550b-a55b \\\n \t--endpoint-type chat --streaming \\\n \t--url localhost:8000 \\\n \t--input-file ./agentic-trace.jsonl \\\n \t--custom-dataset-type mooncake_trace \\\n \t--no-fixed-schedule \\\n \t--concurrency \"$C\"\n done\nExample AIPerf concurrency sweep. Replace the trace, request count, and endpoint details with the workload you want to model.\nStart from the Nemotron 3 Ultra NIM page, accept the governing terms, and select the NIM 2.0.12 tag or the exact published digest. After downloading, find and select a profile. The following lists all profiles packaged in the NIM:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG list-model-profiles\nTo pick an optimized profile for agentic workloads on a four-GPU B200 system, you would set the NIM_MODEL_PROFILE to vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0 and enable speculative decoding:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n export NIM_MODEL_PROFILE=vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -e NIM_SPECDEC_ENABLE=1 \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG\nNVIDIA NIM packages model- and GPU-aware serving optimizations into a deployable microservice. For Nemotron 3 Ultra on a 4xB200 system, these optimizations let organizations serve up to 2.5x more users at 50 TPS/user compared with a NIM-off baseline, while keeping the deployment path practical for real multi-user serving.\nNIM combines validated performance engineering with regular inference-stack updates, CVE handling, hardware validation, and commercial support through NVIDIA AI Enterprise. More performance-optimized NIM configurations are planned across a broader range of models, giving developers additional validated starting points for their own latency, throughput, and cost objectives.\nDownload the Nemotron 3 Ultra NIM 2.0.12 from NGC, run it on your NVIDIA GPU infrastructure, and replay representative traffic with NVIDIA AIPerf to select the Pareto point that meets your application target.\nResources","excerpt":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.","extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 6835 characters.","diagnostics_url":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 6835 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":6835,"summary_length":264,"usable_text_length":6835,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":6835,"summary_length":264}}},"display_formats":["compact","card","full","digest_section","json"]},"daily_stack_record":{"title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","summary":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.","source":"NVIDIA Developer","date":"2026-09-10T17:00:06+00:00","content":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.\nThat tradeoff matters even more for agentic AI workloads, where prompts can be long, context can be reused across steps, and applications often stream extended responses back to users.\nNVIDIA NIM packages model- and GPU-aware serving choices into a deployable microservice. Instead of starting from a blank runtime configuration, developers get a validated serving configuration and a supported deployment path, while retaining the ability to benchmark the NIM against their own traffic.\nInference performance is a system property. Precision and kernels, parallelism, scheduling, batching, memory allocation, prefix reuse, model-specific state caches, and decoding strategy all interact. A configuration is useful only if it improves throughput while staying within the application latency target.\nNIM turns this optimization work into a tested starting point. NVIDIA engineers validate configurations for supported model, GPU, and precision combinations, then package the runtime and model artifacts behind standard APIs. For production deployments, NIM Certified adds regular inference-stack updates, CVE handling, broader hardware validation, and commercial support through NVIDIA AI Enterprise.\nThe benchmark definition used throughout this article is:\n- Hardware: 4xB200\n- Agentic workload: 64K/400/76% KV reuse/50 TPS/user (20 ms ITL)\nThis Pareto chart compares the open-source baseline serving stack (NIM Off) against the fully optimized NIM 2.0.12 serving stack (NIM On).\n| Configuration | Native 256K max context | What it represents | \n|---|---|---|\n| NIM Off baseline | 718 tok/s | No NIM optimizations | \n| NIM On (2.0.12) optimized serving stack | 1,997 tok/s 2.5x v. Baseline | Optimized NIM stack: cache/state reuse, MTP speculative decoding and associated fixes, autotuned kernels, partial-prefix matching, scheduler, batching, memory, and parallelism tuning | \nThe measured gains come from interacting configuration bundles, not independent switches whose percentages can simply be added. The main optimization layers are:\n- Precision and autotuned model-aware kernels. Autotuned mixture-of-experts and Mamba kernels map the hybrid architecture efficiently to NVIDIA Blackwell GPUs.\n- Parallel execution. Tensor parallelism distributes the model across four GPUs, while expert-aware execution improves utilization for the mixture-of-experts layers.\n- Prefix and model-state reuse. Prefix caching avoids recomputing repeated context, partial-prefix matching recovers reuse when only part of a prefix matches, and Mamba state-cache settings are tuned for the model architecture.\n- Scheduler, batching, and memory tuning. Concurrent-sequence limits, batched-token limits, block size, and GPU-memory allocation keep more work in flight without crossing the latency target.\n- MTP speculative decoding. NIM 2.0.12 optimized serving stack (including MTP) adds MTP and its associated fixes to the same optimized serving stack. The incremental benefit depends on acceptance rate and available memory headroom.\nThe published curves are a starting point, not a promise that every application will see the same result. The fastest way to determine fit is to replay representative traffic and build a Pareto curve for the latency metric that matters to your users.\n- Deploy the exact software versions. Use NIM 2.0.12 (or newer version), and pin the image tag or digest for every run.\n- Prepare representative traffic. Use a Mooncake-format JSONL trace or capture controlled NIM requests, with appropriate access controls and sanitization for sensitive data.\n- Measure performance. Use NVIDIA AIPerf to replay representative traffic and get perf benchmarks\n- Select the Pareto point that meets the SLO. Compare output throughput among points that satisfy the SLO/latency constraints and determine fit for deployment:\nfor C in 1 4 8 16 32 64; do\n   aiperf profile \\\n \t--model nvidia/nemotron-3-ultra-550b-a55b \\\n \t--endpoint-type chat --streaming \\\n \t--url localhost:8000 \\\n \t--input-file ./agentic-trace.jsonl \\\n \t--custom-dataset-type mooncake_trace \\\n \t--no-fixed-schedule \\\n \t--concurrency \"$C\"\n done\nExample AIPerf concurrency sweep. Replace the trace, request count, and endpoint details with the workload you want to model.\nStart from the Nemotron 3 Ultra NIM page, accept the governing terms, and select the NIM 2.0.12 tag or the exact published digest. After downloading, find and select a profile. The following lists all profiles packaged in the NIM:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG list-model-profiles\nTo pick an optimized profile for agentic workloads on a four-GPU B200 system, you would set the NIM_MODEL_PROFILE to vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0 and enable speculative decoding:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n export NIM_MODEL_PROFILE=vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -e NIM_SPECDEC_ENABLE=1 \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG\nNVIDIA NIM packages model- and GPU-aware serving optimizations into a deployable microservice. For Nemotron 3 Ultra on a 4xB200 system, these optimizations let organizations serve up to 2.5x more users at 50 TPS/user compared with a NIM-off baseline, while keeping the deployment path practical for real multi-user serving.\nNIM combines validated performance engineering with regular inference-stack updates, CVE handling, hardware validation, and commercial support through NVIDIA AI Enterprise. More performance-optimized NIM configurations are planned across a broader range of models, giving developers additional validated starting points for their own latency, throughput, and cost objectives.\nDownload the Nemotron 3 Ultra NIM 2.0.12 from NGC, run it on your NVIDIA GPU infrastructure, and replay representative traffic with NVIDIA AIPerf to select the Pareto point that meets your application target.\nResources","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 6835 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 6835 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":6835,"summary_length":264,"usable_text_length":6835,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":6835,"summary_length":264}},"tags":[]},"fallback_formats":["markdown","json","html"],"actions":{"read":"/item/79654","export_markdown":"/api/items/79654/export?format=markdown","export_json":"/api/items/79654/export?format=json","diagnose":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/"},"formats":{"full":{"id":79654,"title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","source":"NVIDIA Developer","author":null,"published_at":"2026-09-10T17:00:06+00:00","locale":"en","topic":"ai","tags":[],"excerpt":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.","full_text":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.\nThat tradeoff matters even more for agentic AI workloads, where prompts can be long, context can be reused across steps, and applications often stream extended responses back to users.\nNVIDIA NIM packages model- and GPU-aware serving choices into a deployable microservice. Instead of starting from a blank runtime configuration, developers get a validated serving configuration and a supported deployment path, while retaining the ability to benchmark the NIM against their own traffic.\nInference performance is a system property. Precision and kernels, parallelism, scheduling, batching, memory allocation, prefix reuse, model-specific state caches, and decoding strategy all interact. A configuration is useful only if it improves throughput while staying within the application latency target.\nNIM turns this optimization work into a tested starting point. NVIDIA engineers validate configurations for supported model, GPU, and precision combinations, then package the runtime and model artifacts behind standard APIs. For production deployments, NIM Certified adds regular inference-stack updates, CVE handling, broader hardware validation, and commercial support through NVIDIA AI Enterprise.\nThe benchmark definition used throughout this article is:\n- Hardware: 4xB200\n- Agentic workload: 64K/400/76% KV reuse/50 TPS/user (20 ms ITL)\nThis Pareto chart compares the open-source baseline serving stack (NIM Off) against the fully optimized NIM 2.0.12 serving stack (NIM On).\n| Configuration | Native 256K max context | What it represents | \n|---|---|---|\n| NIM Off baseline | 718 tok/s | No NIM optimizations | \n| NIM On (2.0.12) optimized serving stack | 1,997 tok/s 2.5x v. Baseline | Optimized NIM stack: cache/state reuse, MTP speculative decoding and associated fixes, autotuned kernels, partial-prefix matching, scheduler, batching, memory, and parallelism tuning | \nThe measured gains come from interacting configuration bundles, not independent switches whose percentages can simply be added. The main optimization layers are:\n- Precision and autotuned model-aware kernels. Autotuned mixture-of-experts and Mamba kernels map the hybrid architecture efficiently to NVIDIA Blackwell GPUs.\n- Parallel execution. Tensor parallelism distributes the model across four GPUs, while expert-aware execution improves utilization for the mixture-of-experts layers.\n- Prefix and model-state reuse. Prefix caching avoids recomputing repeated context, partial-prefix matching recovers reuse when only part of a prefix matches, and Mamba state-cache settings are tuned for the model architecture.\n- Scheduler, batching, and memory tuning. Concurrent-sequence limits, batched-token limits, block size, and GPU-memory allocation keep more work in flight without crossing the latency target.\n- MTP speculative decoding. NIM 2.0.12 optimized serving stack (including MTP) adds MTP and its associated fixes to the same optimized serving stack. The incremental benefit depends on acceptance rate and available memory headroom.\nThe published curves are a starting point, not a promise that every application will see the same result. The fastest way to determine fit is to replay representative traffic and build a Pareto curve for the latency metric that matters to your users.\n- Deploy the exact software versions. Use NIM 2.0.12 (or newer version), and pin the image tag or digest for every run.\n- Prepare representative traffic. Use a Mooncake-format JSONL trace or capture controlled NIM requests, with appropriate access controls and sanitization for sensitive data.\n- Measure performance. Use NVIDIA AIPerf to replay representative traffic and get perf benchmarks\n- Select the Pareto point that meets the SLO. Compare output throughput among points that satisfy the SLO/latency constraints and determine fit for deployment:\nfor C in 1 4 8 16 32 64; do\n   aiperf profile \\\n \t--model nvidia/nemotron-3-ultra-550b-a55b \\\n \t--endpoint-type chat --streaming \\\n \t--url localhost:8000 \\\n \t--input-file ./agentic-trace.jsonl \\\n \t--custom-dataset-type mooncake_trace \\\n \t--no-fixed-schedule \\\n \t--concurrency \"$C\"\n done\nExample AIPerf concurrency sweep. Replace the trace, request count, and endpoint details with the workload you want to model.\nStart from the Nemotron 3 Ultra NIM page, accept the governing terms, and select the NIM 2.0.12 tag or the exact published digest. After downloading, find and select a profile. The following lists all profiles packaged in the NIM:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG list-model-profiles\nTo pick an optimized profile for agentic workloads on a four-GPU B200 system, you would set the NIM_MODEL_PROFILE to vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0 and enable speculative decoding:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n export NIM_MODEL_PROFILE=vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -e NIM_SPECDEC_ENABLE=1 \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG\nNVIDIA NIM packages model- and GPU-aware serving optimizations into a deployable microservice. For Nemotron 3 Ultra on a 4xB200 system, these optimizations let organizations serve up to 2.5x more users at 50 TPS/user compared with a NIM-off baseline, while keeping the deployment path practical for real multi-user serving.\nNIM combines validated performance engineering with regular inference-stack updates, CVE handling, hardware validation, and commercial support through NVIDIA AI Enterprise. More performance-optimized NIM configurations are planned across a broader range of models, giving developers additional validated starting points for their own latency, throughput, and cost objectives.\nDownload the Nemotron 3 Ultra NIM 2.0.12 from NGC, run it on your NVIDIA GPU infrastructure, and replay representative traffic with NVIDIA AIPerf to select the Pareto point that meets your application target.\nResources","reading_time_min":5,"extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 6835 characters.","diagnostics_url":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 6835 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":6835,"summary_length":264,"usable_text_length":6835,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":6835,"summary_length":264}}},"quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 6835 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":6835,"summary_length":264,"usable_text_length":6835,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":6835,"summary_length":264}},"actions":{"read":"/item/79654","export_markdown":"/api/items/79654/export?format=markdown","export_json":"/api/items/79654/export?format=json","diagnose":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/"}},"digest":{"id":79654,"title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","source":"NVIDIA Developer","topic":"ai","published_at":"2026-09-10T17:00:06+00:00","excerpt":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.","quality_bucket":"high","quality_reason":"High confidence: full text extraction produced 6835 characters.","reading_time_min":5,"cluster_id":null},"card":{"display_title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","subtitle":"NVIDIA Developer · 2026-09-10","summary":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the…","badges":["quality:high"],"links":{"read":"/item/79654","original":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","diagnose":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/"},"quality_warning":null},"export":{"title":"How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra - NVIDIA Developer","url":"https://developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","summary":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.","source":"NVIDIA Developer","date":"2026-09-10T17:00:06+00:00","content":"Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as possible on available GPU infrastructure while preserving the interactivity that keeps applications responsive.\nThat tradeoff matters even more for agentic AI workloads, where prompts can be long, context can be reused across steps, and applications often stream extended responses back to users.\nNVIDIA NIM packages model- and GPU-aware serving choices into a deployable microservice. Instead of starting from a blank runtime configuration, developers get a validated serving configuration and a supported deployment path, while retaining the ability to benchmark the NIM against their own traffic.\nInference performance is a system property. Precision and kernels, parallelism, scheduling, batching, memory allocation, prefix reuse, model-specific state caches, and decoding strategy all interact. A configuration is useful only if it improves throughput while staying within the application latency target.\nNIM turns this optimization work into a tested starting point. NVIDIA engineers validate configurations for supported model, GPU, and precision combinations, then package the runtime and model artifacts behind standard APIs. For production deployments, NIM Certified adds regular inference-stack updates, CVE handling, broader hardware validation, and commercial support through NVIDIA AI Enterprise.\nThe benchmark definition used throughout this article is:\n- Hardware: 4xB200\n- Agentic workload: 64K/400/76% KV reuse/50 TPS/user (20 ms ITL)\nThis Pareto chart compares the open-source baseline serving stack (NIM Off) against the fully optimized NIM 2.0.12 serving stack (NIM On).\n| Configuration | Native 256K max context | What it represents | \n|---|---|---|\n| NIM Off baseline | 718 tok/s | No NIM optimizations | \n| NIM On (2.0.12) optimized serving stack | 1,997 tok/s 2.5x v. Baseline | Optimized NIM stack: cache/state reuse, MTP speculative decoding and associated fixes, autotuned kernels, partial-prefix matching, scheduler, batching, memory, and parallelism tuning | \nThe measured gains come from interacting configuration bundles, not independent switches whose percentages can simply be added. The main optimization layers are:\n- Precision and autotuned model-aware kernels. Autotuned mixture-of-experts and Mamba kernels map the hybrid architecture efficiently to NVIDIA Blackwell GPUs.\n- Parallel execution. Tensor parallelism distributes the model across four GPUs, while expert-aware execution improves utilization for the mixture-of-experts layers.\n- Prefix and model-state reuse. Prefix caching avoids recomputing repeated context, partial-prefix matching recovers reuse when only part of a prefix matches, and Mamba state-cache settings are tuned for the model architecture.\n- Scheduler, batching, and memory tuning. Concurrent-sequence limits, batched-token limits, block size, and GPU-memory allocation keep more work in flight without crossing the latency target.\n- MTP speculative decoding. NIM 2.0.12 optimized serving stack (including MTP) adds MTP and its associated fixes to the same optimized serving stack. The incremental benefit depends on acceptance rate and available memory headroom.\nThe published curves are a starting point, not a promise that every application will see the same result. The fastest way to determine fit is to replay representative traffic and build a Pareto curve for the latency metric that matters to your users.\n- Deploy the exact software versions. Use NIM 2.0.12 (or newer version), and pin the image tag or digest for every run.\n- Prepare representative traffic. Use a Mooncake-format JSONL trace or capture controlled NIM requests, with appropriate access controls and sanitization for sensitive data.\n- Measure performance. Use NVIDIA AIPerf to replay representative traffic and get perf benchmarks\n- Select the Pareto point that meets the SLO. Compare output throughput among points that satisfy the SLO/latency constraints and determine fit for deployment:\nfor C in 1 4 8 16 32 64; do\n   aiperf profile \\\n \t--model nvidia/nemotron-3-ultra-550b-a55b \\\n \t--endpoint-type chat --streaming \\\n \t--url localhost:8000 \\\n \t--input-file ./agentic-trace.jsonl \\\n \t--custom-dataset-type mooncake_trace \\\n \t--no-fixed-schedule \\\n \t--concurrency \"$C\"\n done\nExample AIPerf concurrency sweep. Replace the trace, request count, and endpoint details with the workload you want to model.\nStart from the Nemotron 3 Ultra NIM page, accept the governing terms, and select the NIM 2.0.12 tag or the exact published digest. After downloading, find and select a profile. The following lists all profiles packaged in the NIM:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG list-model-profiles\nTo pick an optimized profile for agentic workloads on a four-GPU B200 system, you would set the NIM_MODEL_PROFILE to vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0 and enable speculative decoding:\n export NGC_API_KEY=<your-personal-api-key>\n export LOCAL_NIM_CACHE=$HOME/.cache/nim\n export NIM_TAG=2.0.12\n export NIM_MODEL_PROFILE=vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0\n mkdir -p \"$LOCAL_NIM_CACHE\"\n echo \"$NGC_API_KEY\" | docker login nvcr.io \\\n   --username '$oauthtoken' --password-stdin\n docker run --gpus all --shm-size=16GB \\\n   -e NGC_API_KEY \\\n   -e NIM_MODEL_PROFILE \\\n   -e NIM_SPECDEC_ENABLE=1 \\\n   -v \"$LOCAL_NIM_CACHE:/opt/nim/.cache\" \\\n   -p 8000:8000 \\\n   nvcr.io/nim/nvidia/nemotron-3-ultra-550b-a55b:$NIM_TAG\nNVIDIA NIM packages model- and GPU-aware serving optimizations into a deployable microservice. For Nemotron 3 Ultra on a 4xB200 system, these optimizations let organizations serve up to 2.5x more users at 50 TPS/user compared with a NIM-off baseline, while keeping the deployment path practical for real multi-user serving.\nNIM combines validated performance engineering with regular inference-stack updates, CVE handling, hardware validation, and commercial support through NVIDIA AI Enterprise. More performance-optimized NIM configurations are planned across a broader range of models, giving developers additional validated starting points for their own latency, throughput, and cost objectives.\nDownload the Nemotron 3 Ultra NIM 2.0.12 from NGC, run it on your NVIDIA GPU infrastructure, and replay representative traffic with NVIDIA AIPerf to select the Pareto point that meets your application target.\nResources","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//developer.nvidia.com/blog/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra/","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 6835 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 6835 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":6835,"summary_length":264,"usable_text_length":6835,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":6835,"summary_length":264}},"tags":[],"format_contract_version":"news_item_formats.v1"}}}