Nvidia Says AI Agents Must Be Stopped Even When They Make Wrong Judgments, Calls for Stronger Security Boundaries
TECHWORLD ·
✦ AI Summary
Nvidia said that as AI agents spread, security should be viewed not as a problem of individual models but as a control issue across the entire agent stack.
It explained that security controls should be applied across the whole stack, including models, harnesses, and runtime environments, and that clear requirements, accountable owners, testing, and evidence are needed.
It also proposed OpenShell, emphasizing the need for sandbox isolation, access control over data, network, and system resources, and security validation and log preservation.
Nvidia, in response to the spread of AI agents, emphasized strengthening security boundaries and concluded that security should be treated not as a problem of individual models, but as a control issue across the entire agent stack, in the sense that even AI agents' wrong judgments must be stopped. Nvidia laid out an engineering approach to AI security on the 29th.
Nvidia said security controls need to be applied across the entire agent stack, including not only the model but also data, identity, tools, services, and infrastructure. It said this reflects the view that security boundaries must be strengthened across AI systems as a whole.
Nvidia said AI systems need clear security requirements. It also said AI systems need enforceable controls and accountable owners, and stressed the need for testing to prove that security measures work in real-world environments, as well as evidence to demonstrate that those measures are functioning.
Nvidia explained that the rise of the internet and cloud computing has significantly changed the environment in which software operates. Still, it said the basic principles of identity verification, access control, limiting exposure, and verifying security measures remain valid.
AI agents add a new variable. Beyond carrying out fixed commands, agents can reason about situations, use tools, and change their behavior based on the data they encounter. As a result, the object of security design expands from the model itself to the entire environment in which the agent operates.
Nvidia divides the AI agent stack into models, harnesses, and runtime environments. The model's role is to provide capabilities, the harness's role is to configure context, tools, and workflows, and the runtime environment's role is to provide the infrastructure that carries out actual tasks. Because the layers are interconnected, applying security to only one area is not enough.
For example, an agent that updates customer information may come into contact with malicious instructions in an attached document. The agent may attempt to send customer data to an unauthorized external destination. Therefore, the principle of response is not to rely solely on the agent's judgment, and it is necessary to block such data transfers through network policies.
An agent's use of tools and exercise of privileges must be traceable after the fact. To that end, protected logs of tool invocation attempts, authorization decisions, and execution results are needed, and security personnel must be able to trace the tools used and the targets of access attempts.
Privileges need to be managed independently by function. Having permission to modify customer information should not automatically grant permission to send data externally, and while an agent may request additional privileges, self-approval is prohibited.
It is important to set security boundaries in the execution environment itself so that damage can be reduced even if an agent makes a wrong judgment. Even in such cases, access to files, network targets, and processes must be restricted.
Nvidia explained that each agent should be assigned a traceable identity and credentials appropriate to its task scope. It also said the range of information access, the scope of system changes, and tasks requiring human approval must be clearly defined in policy, and that human approval procedures are needed for important tasks and privilege changes.
In agent security, the source and integrity of the tools, skills, and dependencies used by the agent must be verified first. If an incident occurs, protected logs recording tool invocation history, authorization decisions, and execution results must be secured, and these logs are used to reconstruct what happened. In addition, response procedures need to include revoking access privileges and isolating the incident.
Nvidia proposed the open-source secure runtime OpenShell as a way to implement such security boundaries. OpenShell has the characteristic of enforcing policies in an external area that the agent cannot directly change. It also provides functions to isolate agent execution in a sandbox environment and control access to data, networks, and system resources.
Partners in the Open Secure AI Alliance are developing related solutions based on OpenShell. Cisco's DefenseClaw added a governance layer. JFrog integrates with OpenShell to scan and verify agent skills and provides a function to apply policies to skill targets that agents can access.
Nvidia stressed that it is important to verify that security controls work before deploying AI agents into real-world environments. It explained that the scope of verification should go beyond functional testing and check whether attempts to obtain credentials beyond authorized scope are blocked and whether attempts to send sensitive data to unauthorized destinations are blocked. It also said the test targets should include privilege escalation attempts and attempts to interfere with monitoring functions.
Nvidia explained that if there are major changes to models, tools, or workflows, the same security tests need to be repeated. It said this is to confirm that existing controls remain effective. It also said it is necessary to clarify accountability for test results and to link failed tests to actual remediation.
Nvidia said problems found during operations need to be reproduced, investigated, and resolved. It also said discovered issues should be turned into repeatable test items. This, it said, has the effect of allowing teams to confirm whether fixes continue to hold in future updates.
Nvidia introduced CrowdStrike SafeMind and Palo Alto Networks Prisma AIRS as related examples of reviewing defense systems. SafeMind was presented as a way to test defense systems based on repeated attack simulations. Prisma AIRS was linked to continuous red-team testing in response to changes in models and applications.
Nvidia then outlined ways to use AI for security defense. Uses for AI included security incident investigations, vulnerability detection, and verification of fixes. It also explained that open models and closed models support security work in different ways.
Closed models were presented as providing managed functions and services. By contrast, open models were presented as allowing direct inspection of related components, adjustment of strategies, and work within self-controlled infrastructure. Nvidia said one advantage of open models in the event of a security incident is that sensitive evidence can be kept in-house, problems can be reproduced, and fixes can be validated on real systems.
Nvidia cited Capital One's VulnHunter and ReversingLabs' Spectra Assure as examples of AI use in security. It then explained that the effectiveness of AI-based security should not be judged simply by whether AI is used. As evaluation criteria, Nvidia cited reproducible findings, verifiable remediation outcomes, and shorter response times.
Nvidia emphasized that stronger AI security requires open collaboration with the security industry. To that end, it said the details of problems that occur, the controls that were effective, and the methods used to verify fixes should be shared as evidence. Nvidia explained that sharing this evidence allows other organizations to strengthen the security of their own AI systems.
Nvidia said its security research and the Open Secure AI Alliance are the entities supporting such collaboration. It explained that they support collaboration by sharing research findings, tools, and expertise with the security community.
Nvidia explained that security advancement is needed as AI agents take on more capabilities. It said the security approach should shift in a way that reduces dependence on specific models or single solutions, and that enforceable security boundaries must be built across the entire agent stack. It also said clear accountability must be assigned and evidence of actual security controls working must be continuously secured, adding that this would allow security levels to improve alongside the advancement of AI capabilities.
Source: TECHWORLD · Park Gyu-chan
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407483
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.