4AIVN
Back to News

Nvidia NemoClaw: the security platform of OpenClaw for enterprises

Published on 17 March, 2026
Nvidia NemoClaw: the security platform of OpenClaw for enterprises

Quick Summary

Nvidia shocked GTC 2026 by launching NemoClaw, an enterprise-grade AI agent platform addressing OpenClaw's security flaws. Developed alongside Peter Steinberger, NemoClaw uses OpenShell to prioritize privacy and data security. Nvidia also introduced the new Vera CPU with superior performance for agentic AI, expects $1 trillion in AI chip revenue by 2027, and formed the Nemotron Alliance to drive open-source AI, solidifying its leadership in the AI agent era.

Enterprise IT departments surely ban OpenClaw on internal computers. The reason is not that the tool is ineffective, but because nobody can control what company data flows through it. This is a risk that businesses face when deploying AI agents without a reliable security solution. At GTC 2026, Nvidia answered directly with NemoClaw, a platform built on OpenClaw but adding the entire enterprise-grade security layer that the original version lacked.

What is OpenClaw and why enterprises are hesitant to use it?

If you don't know what OpenClaw is, here is the quickest way to understand: instead of sitting and instructing the AI step-by-step, OpenClaw allows you to create autonomous AI agents that work continuously without your intervention. Developed by engineer Peter Steinberger, who has since joined OpenAI, this platform still grows strongly globally, especially in China, even though tech giants like Gemini and Claude have blocked its API connections.

The problem is that OpenClaw was designed for individuals and small teams, not for enterprises with sensitive data. When incorrectly installed or run with default configurations, the AI agent can access and process internal data without any control layers. Governments in many countries and giants like Google and Anthropic have repeatedly issued security warnings about this issue, which is why most enterprises still stay on the sidelines despite knowing the tool's potential. This is exactly the gap that Nvidia saw and decided to fill.

How NemoClaw solves the security puzzle

Instead of building a brand-new agent platform, Nvidia collaborated directly with Peter Steinberger to develop NemoClaw on the existing OpenClaw foundation. CEO Jensen Huang stated at GTC 2026 that every company needs an OpenClaw strategy, and NemoClaw is Nvidia's way of bringing that strategy into reality safely.

The heart of NemoClaw is an open-source execution environment called OpenShell. Imagine it simply: instead of letting the AI agent run freely across the system like an unsupervised new employee, OpenShell locks it in a separate workspace with rules defined by the business itself. Specifically, OpenShell does three main things:

  • Enforces guardrails based on each organization's internal policies, meaning each business decides what the AI agent can and cannot do.
  • Keeps AI models running in a separate sandbox environment, preventing them from accessing data beyond their permitted scope.
  • Adds data privacy protections before any information is processed, while increasing scalability as demand grows.

What do enterprises specifically gain when using NemoClaw?

Three practical benefits NemoClaw brings compared to using OpenClaw, as provided by Nvidia:

  • Data control: The IT department can define exactly which documents and systems the AI agent is allowed to access and what it can do with that data. No more AI agents running wild without anyone knowing what they are reading.
  • Flexible AI model selection: Businesses are not locked into a single vendor. NemoClaw supports Nvidia's NemoTron, Anthropic's Claude, OpenAI's GPT, and any other open AI models, allowing cloud model access right on local devices without relying on specific hardware.
  • No infrastructure changes needed: NemoClaw runs on top of existing OpenClaw setups, meaning teams currently using OpenClaw can upgrade to NemoClaw without starting over.

NemoClaw is currently in the alpha stage, meaning it is still being finalized before the official launch. Currently, NemoClaw has open-sourced its code on GitHub for those who need higher customization. This is a point to note if you are considering enterprise deployment right now.

What else is notable at GTC 2026 besides NemoClaw?

NemoClaw is just one part of Nvidia's massive wave of announcements at GTC 2026. Other key highlights include:

  • Next-generation Vera CPU: Designed specifically for the AI agent era with double the performance and 50% faster speed than traditional CPUs, optimized for complex reinforcement learning tasks.
  • $1 trillion revenue forecast: Nvidia expects revenue from Blackwell and Vera Rubin AI chips to reach this level by 2027, reflecting the company's massive bet on the booming AI agent wave.
  • Nemotron Alliance: An open collaborative initiative to share resources and computing capacity in the open-source AI domain, drawing the participation of many industry giants.
  • Groq 3 and DLSS 5: The Groq 3 language processing unit and DLSS 5 graphics technology were also announced, expanding Nvidia's AI ecosystem beyond agents and into game graphics.

NemoClaw is the bridge bringing AI agents from individuals to enterprises

OpenClaw has proven that AI agents work effectively in practice. The issue is not the technology, but trust—and trust in an enterprise environment comes from control, transparency, and internal policy compliance. NemoClaw does not try to replace OpenClaw, but builds exactly that layer on top of it.

If NemoClaw works as promised when officially released, this could be the key to getting AI agents widely deployed in enterprises, instead of being blocked by IT departments for security reasons. That is precisely the real market Nvidia is targeting.

Discussion (0)

Log in to join the discussion.

No comments yet. Be the first!

Related Articles

Save AI Agent Tokens with Ponytail and Caveman

Token management is a hot topic in the AI community. In summer 2026, two open-source skills for saving tokens are being widely discussed: Ponytail cuts generated code lines by up to 54%, while Caveman slashes agent response tokens by 65%. Both target a familiar pain point for users of AI coding agents like Claude Code, Codex, or Gemini CLI—ballooning token costs—yet solve it from completely different angles: one trims unnecessary code, while the other trims unnecessary words.The Token Waste Problem in the AI Agent EraToday's AI Agents do not just answer single prompts; they operate in autonomous agentic loops by reading files, analyzing project structures, writing code, running builds, and checking for errors. Throughout this process, tokens are wasted mainly across three channels:Over-engineering (Excess Code): Instead of using native language or browser features, agents frequently install extra dependencies or construct unnecessarily complex components.Input Overhead (Excess Context): Build logs, JSON payloads, search results, and instruction files (SKILL.md) consume tens of thousands of input tokens on every API call.Output Bloating (Verbose Prose): Agents explain basic concepts at length before delivering the core answer.Ponytail: Turning your AI Agent into a "Lazy Senior Dev"Created by Dietrich Gebert, Ponytail is designed with a core philosophy: "The best code is the code you never wrote." Ponytail forces an AI Agent to think like a seasoned senior developer who always looks for the simplest, lowest-effort solution.Self-Questioning Ladder Before Writing CodeBefore touching any code, Ponytail requires the agent to run through a self-questioning ladder:YAGNI: Is this feature really necessary? If not, skip it immediately.Reusability: Is there an existing function or component in the codebase?Standard Library: Can the language's standard library handle it?Native Platform: Is it supported natively by the browser or OS? (e.g., using a native <input type="date"> instead of installing a heavy Flatpickr library)Installed Dependencies: Can already installed packages in package.json resolve it?One-liner: Can it be written in a single line of code?Only when the above steps fail will the Agent proceed to write the minimal working code required.Do Benchmark Results Match Reality?In tests using Claude Code (Haiku 4.5) on a full-stack FastAPI + React template, Ponytail reduced lines of code by 54% while maintaining 100% application safety.This figure was published by the author after the community pointed out baseline flaws in the initial benchmark (which claimed 80-94% reductions), so it should be treated as a reference signal rather than an independently verified statistic.Caveman: A Token Compression Ecosystem, Not Just a SkillWhile Ponytail targets generated code, Julius Brussee's Caveman attacks both agent input and output with the catchy slogan: "why use many token when few token do trick". Current Caveman is no longer a single skill, but a multi-layered toolkit.Caveman Proxy: Compressing Input DataA local proxy sits between the Agent and the API provider, automatically routing all traffic. It detects payload types like JSON, error logs, git diffs, or search results, compressing them to keep only essential content needed for the answer, while saving copies on disk for byte-exact recovery when required. In a 54-run pinned benchmark on Claude Code, this mechanism used 33.2% fewer input tokens than direct runs while passing all exact-answer checks.Caveman Skill: Compressing Response LanguageThis part returns agent communication to primal caveman-speak: dropping conversational filler and getting straight to the point without altering code or command integrity. Code, commands, and error logs remain byte-exact, with only explanatory prose compressed.Pixel Mode: Rendering Skill Files to ImagesVerbose SKILL.md files are rendered into PNG images upon skill installation, leveraging vision capabilities of modern LLMs to lower prompt token loads. Measured on Caveman's own skill file, this reduced size from ~1,069 down to an estimated 415 tokens, or 61%.The 61% reduction was measured on a single case (Caveman's own SKILL.md), not as an average across all skill files—actual results will vary based on file length and structure.Caveman Learn: Self-Diagnosing Token BottlenecksThe caveman learn command automatically reads local agent session history (running locally without accounts), scores the current setup, and pinpoints exact token sinks for user remediation.Numbers to Keep in Mind Reading Caveman MarketingCaveman's official documentation includes an "honest number warning": Caveman Skill alone reduces output tokens, while input and reasoning tokens remain mostly unchanged unless the proxy is enabled, adding ~1,000-1,500 input tokens per turn for the skill prompt. The 65% output token, 33.2% input token (via proxy), and 61% skill token (via Pixel Mode) metrics are separate measurements and do not stack into a single combined number—read context carefully before quoting.Can Ponytail and Caveman Be Combined in One Session or Project?Ponytail and Caveman do not conflict. Caveman keeps what the agent reads (input) and says (output) as concise as possible, while Ponytail ensures what the agent writes (code) is strictly minimal. Because their mechanisms do not overlap, they can be used side-by-side in the same session.While no independent benchmark has measured the exact savings of using both simultaneously, combining their individual metrics theoretically reduces overall session tokens significantly—making it well worth testing on your real codebase.Workflow Integration GuideBoth Ponytail and Caveman support quick installation for popular AI coding tools like Claude Code, Codex, Gemini CLI, Cursor, and Windsurf.Installing PonytailFor Claude Code:claude plugin marketplace add DietrichGebert/ponytail && claude plugin install ponytail@ponytailFor other agents without plugin marketplace support, copy rule files directly from the GitHub repository into your project directory.Installing CavemanFor Claude Code:claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@cavemanFor Gemini CLI:gemini extensions install https://github.com/JuliusBrussee/cavemanInstalling for Cursor, Windsurf, Cline, and OthersCaveman can be installed via the unified registry:npx skills add JuliusBrussee/caveman -a <agent-name>Will You Choose Ponytail, Caveman, or Both?Token optimization is not just about saving money; it keeps AI Agent context clean and prevents context drift during long sessions. However, the larger takeaway from both skills is not to blindly accept advertised percentages—even from authors—since each figure typically measures a specific scenario rather than a general average. The surest approach is running benchmarks on your own codebase before adding Ponytail, Caveman, or both to your daily workflow.

Nam
26 Aug, 2026
Automate Excel & Google Sheets Reports with OpenAI Codex

Automating Excel and Google Sheets reporting is no longer exclusive to software engineers. With the rapid evolution of AI models like GPT, office workers can now create custom workflow automation tools using simple instructions with Codex, freeing up hours of repetitive daily tasks. Why Excel Formulas and VBA Are No Longer Enough For weekly recurring reports or automated integrations with email, Slack, and messaging apps, traditional methods like nested Excel formulas or recording VBA (Visual Basic for Applications) macros require specific technical skills and break easily whenever a single column in the source file changes. This is the exact gap that OpenAI Codex fills: you describe precisely what needs to be done in natural language, and Codex generates complete Python or Google Apps Script code ready to run in seconds. Codex Is Not a Single Product A common misconception is how Codex is accessed: it can be used via CLI in the terminal, IDE extensions in VS Code, Codex Web on the cloud at https://chatgpt.com/codex/cloud for developers and coding enthusiasts, as well as desktop applications for both macOS and Windows. For office workers unfamiliar with command-line tools, the easiest way to start is downloading the Codex desktop app. It lets you manage multiple agents simultaneously right from a visual interface without opening a terminal or configuring API keys—just log in with your existing ChatGPT account. What an Example Prompt Looks Like You only need to provide Codex with a detailed prompt like: "Write a Python script that reads the Excel file 'sales_raw.xlsx', filters orders with status 'Completed', calculates total revenue by branch, and exports the result to 'revenue_report.xlsx' with dark blue header styling." Codex will instantly generate standard, high-quality code, and you simply run the script to get your finalized report. Automating Local Excel Reports with Python and Codex For Excel files stored locally, combining Codex with two popular Python libraries—pandas and openpyxl—delivers outstanding processing speed. Pandas handles hundreds of thousands of data rows in seconds, while openpyxl manages cell formatting, header colors, and formula insertion as illustrated in the example above. Automating Google Sheets in the Cloud When your team collaborates on Google Sheets instead of offline Excel files, the Python library gspread or Google Apps Script is the ideal choice. Codex can write code that connects directly to the Google Sheets API via a Service Account (JSON credential file) to read and write data continuously without opening a browser. A Sample Workflow Automatically pull new form submissions from Google Forms into the spreadsheet Automatically categorize customer feedback by priority Automatically send daily summary emails to leadership at 17:00 This entire pipeline runs in the background with zero manual intervention once configured. 4-Step Implementation Process for Non-Coders Step 1 - Standardize input data: Ensure the Excel or Google Sheets file has clean, clear headers without arbitrary merged cells. Step 2 - Write clear prompts for Codex: Explicitly mention file names, column names, filtering/calculation steps, and desired output formats. Step 3 - Test and paste errors for AI auto-debugging: If the script encounters an error, copy the full error traceback and paste it back into Codex for automatic correction. Step 4 - Schedule automatic runs: Use Windows Task Scheduler (Windows) or Cron jobs (macOS/Linux) to run the script automatically on a set schedule. Risks and Considerations Before Handing Reports Over to AI Codex is powerful but not completely free; however, simple operations with Excel or Google Sheets consume very little quota, so light users can comfortably rely on the free tier. If heavier workloads are required, consider Plus or Pro plans—avoid the Go tier as Codex capabilities there offer little advantage over the free plan. Never paste system passwords, financial records, or real customer data into public AI chat interfaces. When asking Codex to generate code, always substitute sensitive details with dummy data of the same structure. For critical validations such as monetary amounts or tax formulas, do not trust AI-generated code blindly on the first run. Always cross-check the output during the first 1-2 executions to ensure the logic perfectly matches your actual business requirements. Automating reports with Codex does not make you a programmer, nor should it. The true value lies in understanding your own data and business processes well enough to describe them clearly to AI—Codex handles the coding. If your company has weekly recurring reports, start today by picking the simplest one, writing a prompt following the sample above, and running your first automated test.

Nam
24 Aug, 2026
Claude Opus 5 Launches, Closing In on Fable 5

Anthropic has launched Claude Opus 5 at the same price as Opus 4.8 while raising response quality close to Fable 5, a model that costs twice as much. In other words, with near-Fable performance at half the price, most users will likely choose Opus 5 as their default and reserve Fable 5 for the small number of tasks that truly require the highest capability ceiling. What upgrades does Claude Opus 5 bring? According to Anthropic's launch announcement, Claude Opus 5 is the most capable Opus model to date and the first Opus release in the Claude 5 generation. Anthropic describes it as proactive and capable of deep reasoning, approaching the highest intelligence of Claude Fable 5 across many domains while using only half the token budget. The API model ID is claude-opus-5. Like Opus 4.8 and Fable 5, it has a default and maximum context window of one million tokens, a 128,000-token output limit, and thinking enabled by default. It has become the default model on Claude Max and the most powerful model available on Claude Pro. It is also offered through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and GitHub Copilot. Why will many users choose Opus 5 over Fable 5? The answer is not limited to price. Four factors make Opus 5 likely to become the default choice for daily work while Fable 5 moves into a specialized role for a small number of exceptional cases. It wins more real-world evaluations than it loses On Frontier-Bench v0.1, Anthropic's automated coding evaluation, Opus 5 scores 43.3% while Fable 5 reaches only 33.7%, a gap of almost ten points in favor of Opus 5. On CursorBench 3.2 at maximum effort, Opus 5 reaches about 70.1%, less than half a percentage point behind Fable 5 while costing only half as much. Across evaluations where both models have published results, Opus 5 wins more often than it loses, and its victories are generally larger than its defeats. The fastest way to verify this is to run the same task on both models at comparable effort levels and compare the output quality instead of relying only on published benchmarks. No mandatory 30-day data retention Fable 5 and Mythos 5 are Covered Models that require prompts and outputs to be retained for 30 days for safety purposes. They do not support zero data retention (ZDR) on any platform, even when an organization already has a ZDR agreement. Opus 5, by contrast, can still operate under ZDR like Opus 4.8. For teams handling legal, medical, or financial data, this difference alone may remove Fable 5 from consideration without any performance comparison. Fewer interruptions from safety filters Anthropic says the cybersecurity classifier intervenes about 85% less often with Opus 5 than with Fable 5. For coding agents that run for hours or overnight, a request being blocked midway because it touches a safety threshold is a real workflow risk, and Opus 5 significantly reduces that frequency. Adjustable effort makes budgets easier to predict Opus 5 supports adaptive thinking with effort ranging from low to maximum. Low or medium works for fast responses and high-volume workloads, while high or maximum suits complex coding, deep research, and multi-step workflows. Because teams pay according to the selected effort instead of being locked into a fixed Fable 5 cost level, they can optimize the budget for each task rather than paying the highest rate on every request. Initial impressions after trying Opus 5 After using Opus 5 for daily writing and coding work, the clearest impression is that it is substantially smarter than Opus 4.8, especially in understanding intent on the first request without repeated explanation. For tasks such as summarizing long documents, writing code with complex branching logic, or preparing a multi-step plan, Opus 5 works smoothly and loses the thread less often than the earlier version. There is still a gap compared with Fable 5, although it is smaller than expected. On work that demands deep reasoning or autonomous execution across many consecutive steps without intervention, Fable 5 remains slightly more dependable and makes fewer mistakes. For most daily work, however, that difference is difficult to notice without placing both models side by side. If you are using Opus 4.8, this is a sensible time to upgrade. If you are choosing between Opus 5 and Fable 5 for ordinary work, Opus 5 is almost certainly sufficient without paying the premium. When is Fable 5 still the right choice? Fable 5 retains an advantage on the hardest work. On SWE-bench Pro, which uses real GitHub issues and is considered one of the strictest measures of practical coding, Fable 5 scores about 80% while Opus 5 reaches roughly 79%, a small gap that still favors Fable. Fable 5 is also the only model Anthropic positions in the Mythos class, meaning its overall capability is designed to exceed Opus. This distinction is clearest in specialized fields such as expert medical analysis and autonomous research that continues for days without supervision. In other words, Opus 5 wins in daily coding and knowledge work, while Fable 5 retains its edge on the hardest problems and fields requiring the highest possible reliability. For most users and small teams, those problems represent a small portion of daily work, making the twofold price difference difficult to justify unless their workload falls directly into that category. Quick comparison: Opus 5 vs. Fable 5 CriterionClaude Opus 5Claude Fable 5 Input price$5/million tokens$10/million tokens Output price$25/million tokens$50/million tokens Context1 million tokens1 million tokens Maximum output128,000 tokens128,000 tokens Frontier-Bench v0.1 (coding agent)43.3%33.7% SWE-bench Pro (practical coding)~79%~80% Data retentionSupports zero data retentionMandatory 30-day retention, no ZDR Safety-filter interventionAbout 85% lowerHigher Best fitDaily work, coding agents, sensitive dataDifficult research, multi-day autonomous projects, specialized medical analysis Can Opus 5 really compete with GPT-5.6? On paper, the answer is yes, but not across every category. Opus 5 leads GPT-5.6 Sol in reasoning about novel situations, computer use, and most public coding evaluations, while GPT-5.6 Sol remains ahead on some command-line and information-retrieval tests. Neither wins outright, but for the first time a mid-priced Anthropic model stands level with, and in several areas ahead of, OpenAI's flagship model. The more useful question is not which model is stronger overall but which one fits your work. If daily tasks center on code, long documents, and multi-step execution, Opus 5 is a compelling choice on both price and quality. If you already rely on the OpenAI ecosystem or need a specific GPT-5.6 strength, the switching cost may not be worthwhile. The most reliable answer is still to run the same job on both models, because benchmark tables do not always reflect real experience.

Nam
25 Jul, 2026
Gemini 3.6 Flash Launches but Disappoints in Practice

Google announced Gemini 3.6 Flash on July 21, 2026, with sharp benchmark gains over 3.5 Flash: DeepSWE rose from 37% to 49%, MLE Bench from 49.7% to 63.9%, and OSWorld Verified reached 83%. Yet 4AIVN's hands-on experience tells a very different story. The model handles small jobs reasonably well, but a multi-step plan can make it forget the objective, skip steps, and drift halfway through the work. Stronger benchmarks do not reflect real-world use According to Google's official announcement, Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, while tests such as DeepSWE show token reductions of up to 65%. Its input window reaches 1,048,576 tokens and its output limit is 65,536 tokens, impressive numbers on paper. The problem is that these figures come from designed tests with a fixed objective and a relatively contained run. That is not how a real plan operates. Production work changes continuously in response to feedback rather than ending after one self-contained attempt. Following a long plan is the critical weakness In hands-on use, Gemini 3.6 Flash performs poorly as soon as it moves beyond a single task. Give it a small job with explicit checks and it can work well with few unnecessary loops. Give it a multi-step plan and it may forget the original objective, skip previously agreed steps, or drift after several turns. When corrected, it sometimes apologizes and then repeats the same mistake instead of actually fixing it. A one million token window describes input capacity, not memory quality. The model may be able to “see” the full context and still miss details during execution; one overlooked constraint can push the entire plan off course. This is not a rare random failure but a repeated weakness that is difficult to ignore. Gemini 3.6 Flash is strong at completing one job quickly, but it is not yet dependable at completing a sequence of jobs correctly. That is the gap the benchmarks do not measure. A 17% price cut may not match the quality Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, about 17% below the $9 output price of 3.5 Flash. On the surface, this is a sensible improvement: lower cost and higher benchmark scores. But if long tasks are executed poorly, the savings can quickly disappear through repeated reminders, corrections, and complete reruns of the plan. Gemini 3.5 Flash Lite is cheaper still at $0.30 per million input tokens and $2.50 per million output tokens, but it targets simple classification and data transformation workloads that do not require the model to preserve a long plan. What do you gain and lose with Gemini 3.6 Flash? Objectively, this is not a failed upgrade. Google has likely made careful tradeoffs among output quality, speed, and cost, even if real-world behavior does not fully meet the high expectations attached to its engineering team. The improvements are real rather than purely theoretical: responses are faster, output costs are lower, and the model is efficient on short, narrow tasks such as content classification, writing one code function, or answering a specific question. In those cases, it keeps unnecessary loops to a minimum. The cost becomes visible when work extends beyond a few steps. The more constraints and earlier decisions the model must preserve, the more likely it is to drift. For coding agents or long workflows already running reliably on Claude Fable 5 or GPT 5.6, there is not yet a convincing reason to switch to Gemini 3.6 Flash solely because of benchmarks or lower pricing. Gemini 3.5 Pro is still the model to wait for Google says Gemini 3.5 Pro is still being tested with partners and will be released broadly when it is ready. The central story of this launch is therefore the sizeable gap between benchmarks and real work. Anyone looking for a dependable agent for long-running workflows may still need to wait and see whether 3.5 Pro delivers a genuine step forward. If future releases remain underwhelming in practice, Google risks surrendering its advantage to competitors including Anthropic, OpenAI, and Meta.

Nam
23 Jul, 2026