Skip to content

AI Factory Operations Agent

Deploy specialized AI agents to investigate cluster issues and streamline governed AI factory operations.

Overview

Note

This is an experimental release. Please reference SUPPORT.md prior to installation.

AI Factory Operations Agent is an NVIDIA blueprint for deploying an extensible agentic operations framework built with NemoClaw. Specialized agents gather and correlate evidence across cluster systems, generate clear root cause summaries, and recommend next steps through governed, auditable workflows that teams can adapt to their own AI factory environments.

Supported workflows include:

  • Read-only Kubernetes workload and cluster inspection.
  • Prometheus queries and Grafana dashboard creation.
  • Slurm job root-cause analysis from scheduler and log evidence.
  • Optional cluster-management integrations.
  • Sandboxed command execution with an auditable agent workflow.

Architecture

AI Factory Operations Agent architecture

OpenShell isolates general agent command execution. Kubernetes, observability, Slurm, and Base Command Manager access use module-specific server-side tools and the credentials explicitly configured for those modules. Browser and MCP clients do not receive cluster credentials. Kubernetes Secrets remain mounted only in the pods that require them.

Requirements

  • OS/architecture: a Kubernetes environment compatible with the container images selected in Helm values.
  • Runtime: Kubernetes, kubectl, and Helm 3.
  • LLM: an OpenAI-compatible endpoint or a cluster capable of running chart-managed vLLM.
  • GPU/driver: required only when deploying chart-managed vLLM; requirements depend on the selected model profile.
  • Optional services: Prometheus, Grafana, Slurm, or cluster-management endpoints for the corresponding modules.

Getting Started

In the early access stage, the published Helm chart and container images require an early access invitation to the private NGC registry; nvcr.io login alone does not grant pull access. To request access before installing, complete the onboarding form or email ai-factory-operations-agent@nvidia.com. You can also contact your NVIDIA account team.

Use the complete installation guide for your environment:

Usage

After installation, forward the AI Factory Operations Agent UI service:

kubectl -n mosaic port-forward svc/mosaic-ui 3000:3000

Open http://localhost:3000.

Choose the interface that matches the caller:

Interface Intended caller Contract
Browser UI Cluster operators Interactive chat, dashboards, and enabled operational workflows.
HTTP API Services and scripts POST /api/headless/chat with prompt and sessionKey; returns the completed assistant turn and run metadata.
mosaic CLI Shell and Slurm automation Sends one prompt to an AI Factory Operations Agent URL and waits for the completed assistant turn.
mosaic-mcp Claude Code, Codex, and other MCP clients Local stdio bridge to the AI Factory Operations Agent HTTP service; exposes chat, history, commands, and enabled module tools.

The AI Factory Operations Agent headless skill gives coding agents the HTTP, CLI, and MCP contracts. Install it for Codex with:

install -d ~/.codex/skills/ai-factory-operations-agent-headless
install -m 0644 docs/skills/ai-factory-operations-agent-headless/SKILL.md \
  ~/.codex/skills/ai-factory-operations-agent-headless/SKILL.md

For Claude Code:

install -d ~/.claude/skills/ai-factory-operations-agent-headless
install -m 0644 docs/skills/ai-factory-operations-agent-headless/SKILL.md \
  ~/.claude/skills/ai-factory-operations-agent-headless/SKILL.md

Invoke the installed skill when an agent needs to discover an AI Factory Operations Agent service, ask an operational question, or configure the MCP bridge.

Releases & Roadmap

Contribution Guidelines

Governance & Maintainers

Security

  • Vulnerability disclosure: SECURITY.md
  • Do not file public issues for security reports.

Support

Community

Use GitHub issues and pull requests for project discussions and collaboration. Participation is governed by the project Code of Conduct.

References

License

This project is licensed under the Apache License 2.0.

NVIDIA and third-party attributions and license texts are provided in NOTICE and LICENSE-3rd-party.txt.

About

Deploy AI agents to investigate cluster issues and streamline governed AI factory operations.

Resources

Code of conduct

Contributing

Security policy

Stars

19 stars

Watchers

2 watching

Forks

Contributors

Languages