As businesses move from experimenting with artificial intelligence (AI) to deploying it at scale, Red Hat has introduced Red Hat AI 3.5, bringing new safety, observability and operational capabilities to its enterprise AI portfolio.
The latest release is designed to help IT and platform engineering teams manage AI with the same operational discipline applied to mission-critical infrastructure.
Red Hat said that while many organisations have successfully moved through AI pilots and early deployments, scaling these workloads across the business introduces new challenges around safety, resource management, governance and performance.
Red Hat AI 3.5 aims to address these challenges by providing a foundation for controlling, securing and observing AI workloads across hybrid cloud environments.
Verifying AI Safety Before Deployment
A major addition in Red Hat AI 3.5 is EvalHub, which enables organisations to evaluate AI models before they are deployed.
The capability is designed to support risk-focused safety benchmarking and the creation of regulatory compliance certifications, helping organisations assess potential risks before models, retrieval-augmented generation (RAG) systems and AI agents are put into production.
Evaluated models in the Red Hat AI Catalog also feature built-in Garak benchmark scores, while more than 20 new validated models have been added to the catalogue.
These include models from Google, NVIDIA and Alibaba Cloud, among others.
The validated models now include integrated scores covering AI safety, personally identifiable information (PII) exposure and toxicity risks, giving enterprises greater visibility when assessing potential model risks.
A selection of validated models has also been tagged for tool-calling, providing additional guidance for organisations selecting models for agentic AI applications.
Greater Visibility Into AI Performance
Red Hat AI 3.5 introduces new observability dashboards designed to give platform teams real-time insight into AI infrastructure and application performance.
The dashboards provide metrics covering inference health, GPU utilisation and AI model performance.
Non-administrator users can also access dashboards showing per-user token consumption, supporting greater transparency around AI usage and cost attribution.
For organisations operating AI as a shared service, these capabilities can help teams understand how resources are being consumed across users, models and workloads.
The release also adds MLflow visual agentic tracing, providing greater visibility into the behaviour and performance of AI agents.
Managing Shared GPU Infrastructure
As AI workloads place growing demands on GPU infrastructure, Red Hat AI 3.5 expands its multi-tenancy capabilities to help organisations manage shared computing resources.
Fair-share GPU scheduling allows organisations to manage resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing.
This allows real-time inference workloads to receive priority while background workloads can make use of available capacity.
For organisations requiring stronger isolation between tenants, Red Hat AI now officially supports hosted control planes on Red Hat OpenShift Virtualization.
Each tenant can have a dedicated cluster control plane while the underlying hardware is consolidated. AI workloads running in OpenShift Virtualization virtual machines can also benefit from VM-level isolation across shared, GPU-enabled infrastructure.
Red Hat said these capabilities allow infrastructure providers to operate and upgrade the underlying environment from a single point of control.
Building More Governed AI Agents
Red Hat AI 3.5 also expands capabilities for organisations developing and deploying agentic AI.
Its AutoRAG technology connects enterprise data repositories with agentic applications while adding support for multilingual documents, conversational testing and contextual retrieval.
A visual pipeline allows teams to test and assess RAG configurations before deployment, while pre-configured agent templates provide starting points for common enterprise applications.
These include code review, document processing and research workflows.
The AI Hub also introduces agent development kits and starter kits that integrate frameworks, tools and deployment configurations, allowing agents to operate within sandboxed environments while retaining operational controls and security policies from development through production.
Controlling AI Costs Through Dynamic Computing
Red Hat AI 3.5 introduces Inference-Time Scaling (ITS) to help organisations manage computing resources more efficiently.
The capability dynamically adjusts compute usage according to the complexity of a query, allowing organisations to allocate more resources when more complex reasoning is required and reduce usage for simpler requests.
The release also includes CPU offloading for more efficient GPU memory management, with storage offloading available as a developer preview.
These capabilities are designed to allow models to handle longer conversations and larger documents without necessarily requiring additional GPU hardware.
Red Hat AI 3.5 also supports multimodal serving through vLLM Omni in early access, allowing text, audio and image generation to be served through a unified serving layer.
Expanding AI Across Hybrid and Multi-Cloud Environments
The latest release extends distributed inference capabilities beyond OpenShift to third-party Kubernetes services.
Red Hat said this provides a more consistent model-serving experience across cloud environments, with the technology now generally available on CoreWeave CKS and Microsoft Azure, while Amazon EKS is available as a technology preview.
The platform also expands GPU-as-a-Service and multi-tenancy capabilities, including namespace isolation, hosted control plane support and real-time dashboard visibility into hardware inventories, active utilisation and dynamic GPU capacity borrowing.
For enterprises operating AI across complex infrastructure environments, these capabilities are intended to provide greater consistency and control.
Red Hat AI 3.5 Targets Accountable Enterprise AI
Joe Fernandes, Vice President and General Manager of the AI Business Unit at Red Hat, said the conversation around enterprise AI has moved beyond simply putting AI into production.
“The conversation has moved from getting AI into production to running it at scale as trusted enterprise infrastructure, which requires safety evidence, governed agents, cost attribution and multi-tenancy,” he said.
Fernandes said Red Hat AI 3.5 provides the operational controls, verifiable trust and agentic foundations needed to run AI as a safe and accountable enterprise architecture across hybrid cloud environments.
By bringing AI safety evaluation, observability, multi-tenancy, agent development and resource controls into a unified platform, Red Hat aims to help organisations transition from isolated AI experiments towards governed AI services that can operate across the enterprise.
Red Hat AI 3.5 Now Generally Available
Red Hat AI 3.5 is now generally available and is also available as part of Red Hat AI Factory with NVIDIA.
The release represents Red Hat’s latest effort to give enterprises the infrastructure and governance capabilities needed to scale AI while maintaining visibility over safety, performance, resource utilisation and operational control.


