- Nvidia is reportedly building a trillion-parameter Nemotron 4.
- Nvidia is expanding beyond AI chips into models and software.
Nvidia is developing a new generation of Nemotron AI models, with the largest version expected to contain at least one trillion parameters, according to The Information. The reported project would extend the chipmaker’s work on open models alongside its established GPU and AI computing businesses.
The Nemotron 4 family remains in training, and Nvidia has not announced a release date. Employees working on the project told The Information that the model could be ready as early as late autumn, although Nvidia has not confirmed either the parameter count or the proposed timing.
Nvidia has previously acknowledged that Nemotron 4 is under development. “Nvidia is investing in Nemotron because we believe every company and every country needs accessible frontier open models to strengthen safety and security, accelerate innovation, and provide a foundation they can rely on from one generation to the next,” Kari Briski, Nvidia’s vice-president of generative AI, said.
Nvidia builds beyond the model layer
Nemotron is part of a broader Nvidia AI software portfolio that includes tools for model development, customisation, deployment, and workload routing. The company’s NeMo framework supports stages including training, fine-tuning, and reinforcement learning.
The reported Nemotron 4 project follows several releases that have expanded Nvidia’s position in model development. Nvidia’s current Nemotron family includes models designed for reasoning, coding, tool use, retrieval, and agent-based workloads.
Nvidia is also releasing training resources alongside the models. The company says its open Nemotron datasets contain more than 10 trillion tokens and 40 million post-training samples, alongside training recipes and technical reports.
Nemotron can also be deployed through inference frameworks including vLLM, SGLang, Ollama, and llama.cpp. This gives developers options to run the models across cloud, data-centre, workstation, and edge environments rather than relying solely on Nvidia-hosted services.
The company added another model to the family on 11 August with Nemotron 3.5 Lightning. Reuters said the model is aimed at workloads including code review, tool use, security alert monitoring, and responding to billing questions.
Alongside the model, Nvidia released NeMo Switchyard, software designed to determine which model handles individual parts of an AI workflow. It can route requests between specialised and frontier models rather than sending every task to the same system.
The routing layer is not limited to Nvidia models. Nvidia says Switchyard can work across models from different providers while separating routing decisions from the underlying provider endpoint and model identifier.
Applications can therefore switch models or providers without rewriting the routing integration. Switchyard can also retain context between stages of an agent session when required by the routing policy.
Nvidia has positioned the software partly around inference costs, with routing policies able to assign simpler tasks to smaller models while escalating more difficult requests to larger ones.
LangChain tested the approach using 145 multi-turn agent tasks covering customer support, incident investigation, messaging, issue tracking, email, retrieval, tool use, and long-context work. Requests were routed between Nvidia’s Nemotron 3.5 Lightning and Anthropic’s Claude Opus 4.8.
According to Nvidia, across five LangChain evaluation runs, 7% of calls were sent to Claude Opus 4.8, while the remaining requests were handled by Nemotron 3.5 Lightning. Nvidia reported a 74% cost reduction compared with the frontier-only baseline, alongside an accuracy reduction of about six percentage points.
Nvidia also provides NIM, a set of containers for self-hosting GPU-accelerated inference services for pretrained and customised AI models. The software exposes standard APIs for integration with applications and development frameworks.
The service is not limited to Nemotron. Nvidia says NIM supports thousands of open models through inference frameworks including TensorRT-LLM, vLLM, and SGLang, while optimising inference for combinations of foundation models and Nvidia GPUs.
Its software portfolio extends into application development through AI Blueprints, which provide reference architectures combining products such as Nemotron, NeMo, and NIM for uses including enterprise retrieval and agent workflows.
Open models sit alongside Nvidia’s hardware business
Nvidia’s model development is taking place alongside its role as an infrastructure supplier to companies running models developed elsewhere. IBM and Together AI announced a $240 million multi-year agreement on 11 August to build an AI inference cluster on IBM Cloud using Nvidia systems.
The cluster will use Nvidia HGX B300 systems containing Blackwell 300 processors and Spectrum-X Ethernet networking equipment. The initial deployment is expected to include around 2,000 Blackwell 300 chips, according to Together AI.
Together AI provides infrastructure for training and running open models including DeepSeek, MiniMax, and Kimi. The deal provides a current example of Nvidia hardware being used to run models from other developers while the company expands its own Nemotron family.
That model-neutral approach also appears in Switchyard, which can route workloads across Nvidia models and systems from other providers. Nvidia’s portfolio now spans AI computing systems, Nemotron models, NIM inference software, and routing tools that work across multiple model providers.
Nvidia has also backed open-model development in policy discussions. In July, the company joined Microsoft, Meta, IBM, and other organisations in calling on US lawmakers to avoid broad restrictions on open AI models.
“We should avoid premature restrictions on open models that stifle competition or drive innovation overseas,” Nvidia chief executive Jensen Huang said.
The letter argued that open-weight models allow researchers and developers to examine model behaviour, identify vulnerabilities, and develop safeguards. It called for concerns around technology theft to be addressed through targeted legal and commercial measures rather than broad restrictions on open models.
Nvidia has separately joined an industry coalition focused on developing and sharing AI safety and cybersecurity tools. Reuters reported on the initiative in July after security incidents involving autonomous AI agents.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.
Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
The post Nvidia reportedly builds 1-trillion-parameter Nemotron 4 AI model appeared first on TechWire Asia.
