# AceTeam.ai - Full Documentation > Concatenated corpus of the AceTeam documentation for AI ingestion. > See https://aceteam.ai/llms.txt for the concise index and MCP setup. ## Index ### Agent Setup - [Agent Index (this page)](https://aceteam.ai/docs/agent-index): AI-agent-optimized entry point: MCP install, setup steps, and the full docs index. - [llms.txt](https://aceteam.ai/llms.txt): Machine-readable index of AceTeam docs and MCP setup, served as text/plain. - [MCP Setup](https://aceteam.ai/mcp): One-command setup page for connecting AceTeam to an AI client. - [Create an API Key](https://aceteam.ai/api-keys): Generate an `act_` API key to authenticate the MCP server (shown once). ### Getting Started - [Platform Overview](https://aceteam.ai/docs/platform-overview): What AceTeam is, its key features, and who it's built for. - [Account Setup](https://aceteam.ai/docs/account-setup): Sign up, create your organization, invite team members, and configure billing. - [Organizations Across Your Devices](https://aceteam.ai/docs/organizations-and-devices): How your active organization is remembered separately on each device or browser you sign in on. - [Projects](https://aceteam.ai/docs/projects): Organize agents, workflows, and resources into projects. - [Creating Your First Agent](https://aceteam.ai/docs/first-agent): Build and test your first AI agent in under 5 minutes. - [Creating Your First Workflow](https://aceteam.ai/docs/first-workflow): Build a multi-step automation with the visual workflow editor. - [Connect AceTeam to Claude, Codex, or Claude Code (MCP)](https://aceteam.ai/docs/connect-mcp): Connect the AceTeam MCP server to Claude desktop, Codex, or Claude Code on Windows, Mac, or Linux. The connect URL, OAuth sign-in, and fixes when no tools show. - [Connect AceTeam to Claude Desktop on Windows and Mac](https://aceteam.ai/docs/connect-claude-desktop): Add AceTeam as a custom connector in the Claude desktop app on Windows or Mac: Manage connectors, paste the MCP URL, sign in with OAuth, restart, verify tools. - [Connect AceTeam to Codex (CLI and app) on Windows, Mac, Linux](https://aceteam.ai/docs/connect-codex): Add the AceTeam MCP server to OpenAI Codex CLI or the Codex app on Windows, Mac, or Linux: one command, OAuth login, config.toml, restart, and troubleshooting. - [Connect AceTeam to Claude Code CLI on Linux, Mac, Windows](https://aceteam.ai/docs/connect-claude-code): Add the AceTeam MCP server to Claude Code with one command, sign in with OAuth from the /mcp menu, and verify tools on Linux, Mac, or Windows. - [Ace CLI Setup](https://aceteam.ai/docs/ace-cli-setup): Install the Ace CLI and run AI workflows locally from your terminal. ### Agents - [Agent Builder](https://aceteam.ai/docs/agent-builder): Configure AI agents with the Workbench: models, prompts, parameters, and tools. - [Agent Tools & MCP](https://aceteam.ai/docs/agent-tools-mcp): Extend agents with tools using the Model Context Protocol. - [Onboarding External Harnesses](https://aceteam.ai/docs/external-harness-onboarding): Point Codex, opencode, Cline, or a custom client at the AceTeam MCP and it self-hydrates identity, memory, and the playbook. - [Voice Agents](https://aceteam.ai/docs/voice-agents): Enable real-time voice conversations and phone calls with your agents. - [Ground Agents with Knowledge Collections (RAG)](https://aceteam.ai/docs/knowledge-collections): Build a knowledge base with collections, upload documents and files, link a collection to an agent, and let it retrieve and cite that knowledge (RAG) during chat. - [Deploying Agents to Slack, WhatsApp, and Telegram](https://aceteam.ai/docs/channels-messaging): Connect an agent to Slack, WhatsApp, Telegram, Discord, email, and SMS or phone. Set up a WhatsApp bot, a Telegram bot, and deploy an agent to messaging channels so it replies. - [Giving Agents Persistent Memory](https://aceteam.ai/docs/agent-memory): How agent memory works: store and recall facts across sessions, org-scoped memory files, automatic recall, pinning, and inspecting what an agent remembers. - [Generate Images and Video on Your Own Fabric](https://aceteam.ai/docs/image-video-generation): Render images and short video clips from a text prompt on your own Citadel node, with an optional bring-your-own-key cloud fallback for images. - [Author and Install Reusable Agent Skills](https://aceteam.ai/docs/skills): Package instructions and an optional bundled Flow into a versioned, installable skill, then pin or track its updates per agent. - [Steer a Browser on a Citadel Node with Human Handoff](https://aceteam.ai/docs/cobrowse): Drive a headed browser session on your own hardware from an agent, hand control to a human for login or 2FA, and audit every scripted action. ### Workflows - [Workflow Editor](https://aceteam.ai/docs/workflow-editor): Navigate the visual DAG editor to build multi-step automations. - [Workflow Nodes](https://aceteam.ai/docs/workflow-nodes): Understand the node types available in the workflow editor. - [Workflow Triggers](https://aceteam.ai/docs/workflow-triggers): Start workflows manually, on a schedule, via webhook, or from events. - [Workflow Execution](https://aceteam.ai/docs/workflow-execution): Run workflows, monitor progress, and handle errors. - [Human Approval Steps in Workflows](https://aceteam.ai/docs/human-approvals): Add a human-in-the-loop approval gate to a workflow, then review, approve, or deny pending items from the review queue or over MCP. - [Schedule Shell Commands, Workflows, Agent Prompts, and Reminders](https://aceteam.ai/docs/scheduled-jobs-reminders): Create recurring or one-shot scheduled jobs that run a shell command on your own node, trigger a workflow, run an inline agent prompt, or deliver a static reminder, and inspect each run's history. - [Connect a Database, Publish a Flow as an Endpoint, and Receive Webhooks](https://aceteam.ai/docs/database-connections-endpoints): Give a Flow a database credential to dial directly, publish it as a callable HTTP endpoint or a public run receipt, and register a webhook to receive platform events. ### Platform - [Billing & Credits](https://aceteam.ai/docs/billing): How AceTeam's credit-based billing works. Understand tiers, credits, usage tracking, and how to manage your account. - [Usage Dashboard](https://aceteam.ai/docs/usage-dashboard): Track every dollar your organization spends across LLM calls, hosted instances, storage, and egress, broken down by container, source, and the user, key, schedule, instance, or webhook that drove the cost. - [Adoption Ladder](https://aceteam.ai/docs/adoption-ladder): A behavioural, five-rung readiness ladder for your organization, plus usefulness dimensions beside it, reported by the org_adoption_report MCP tool. - [Publishing Pages and Public Docs](https://aceteam.ai/docs/pages-publishing): Create and publish a page to a shareable /p/ link, control who can view it (public, unlisted, org, private), manage org pages, attach files, and export to PDF. - [Documents, Templates, and E-Signatures](https://aceteam.ai/docs/documents-esignature): Create a reusable document template, generate a document from it, send it for review, and request an e-signature. Covers merge fields, request signature links, signature status, and retention controls. - [Import a Node Folder into Drive and Publish a Public Link](https://aceteam.ai/docs/drive-import-publish): Pull a folder from your own Citadel node into AceTeam Drive with its subfolder tree intact, then mint one public browse and download link over the result. - [Manage Google Calendar Events and Bookable Scheduling Links](https://aceteam.ai/docs/calendar-booking): Read, create, update, and search Google Calendar events, find free time, and (once enabled) publish a Calendly style booking link backed by your own calendar. - [Parse PDFs and Scanned Documents with Sovereign OCR](https://aceteam.ai/docs/ocr-document-parsing): Turn an uploaded PDF or image into structured text on your own fabric, one file or a whole directory at a time, with no central OCR fallback. - [Store Structured Agent Records Without a Custom Database](https://aceteam.ai/docs/records): Write, query, and list small JSON records under a namespace and key, with secret scrubbing on write and scoped access to another user's data. - [Publish a Hosted HTML App at a Public Short Link](https://aceteam.ai/docs/hosted-apps): Create, update in place, and delete a self-contained HTML app served at a public short code, with the same visibility rules as hosted pages. - [Route Chat by Team and Post to Channels with Team Chat](https://aceteam.ai/docs/teams-team-chat): Group members into teams for chat routing, then create channels, open conversations, and post messages on AceTeam's native Team Chat surface. - [Keep a Shared Daily Journal for Humans and Agents](https://aceteam.ai/docs/journal): Append dated, sectioned entries to a per-user journal from an agent or a human, read them back in one canonical format, search across days, and share a day or a single section. - [Push Curated Cards to the Home Feed and Manage What Surfaces](https://aceteam.ai/docs/feed): Surface a report or a suggestion on your reading-list feed, opt a hosted page out of the feed permanently, and read the same assembled feed the web and iOS apps render. ### Sovereign Compute - [AceTeam Executive Overview](https://aceteam.ai/docs/executive-overview): A business-oriented introduction to AceTeam: AI infrastructure you own and control. - [Sovereign AI Compute Platform](https://aceteam.ai/docs/whitepaper): Technical whitepaper: how AceTeam turns distributed hardware into a managed AI factory. - [Fabric Overview](https://aceteam.ai/docs/fabric-overview): Sovereign AI on hardware you own: run open-weight models and self-hosted AI agents locally, with governance and an audit trail built in. - [Citadel Setup](https://aceteam.ai/docs/citadel-setup): Install the Citadel CLI, authenticate, and register your first compute node. - [Coding Agent & File Browser](https://aceteam.ai/docs/coding-agent): Let your AI agent read, edit, and search code on your own machine. Browse files from the web UI. - [Connecting Nodes](https://aceteam.ai/docs/connecting-nodes): Add compute nodes, manage preauth keys, and configure resource routing. - [GPU Compute](https://aceteam.ai/docs/gpu-compute): Deploy models, monitor GPUs, benchmark performance, and manage power on your own hardware. - [Sandboxes](https://aceteam.ai/docs/sandboxes): Run ephemeral Docker containers on your own hardware. Execute code, checkpoint state, and build custom environments. - [ACET Compute Tokens](https://aceteam.ai/docs/acet-tokens): Purchase ACET tokens to pay for GPU inference on the Sovereign Compute Fabric. Programmatic registration for agents. - [Citadel OS](https://aceteam.ai/docs/citadel-os): Pre-built VM image with GPU drivers, Docker, and Citadel pre-installed. Deploy a GPU node in minutes. - [Share Fabric Nodes, List Them on the Marketplace, and Track Earnings](https://aceteam.ai/docs/fabric-marketplace-node-sharing): Grant another organization access to a compute node, list a verified node on the AceTeam marketplace, set your own pricing, and track earnings and withdrawals. ### Safety & Accountability - [AEP Safety Proxy](https://aceteam.ai/docs/safety-proxy): A reverse proxy that intercepts every LLM call, detects PII and toxic content, enforces PASS/FLAG/BLOCK decisions, and tracks cost. Zero code changes. - [Safety Detectors](https://aceteam.ai/docs/safety-detectors): How AEP's pluggable detector architecture works. Built-in detectors for PII, toxicity, and cost anomalies, plus how to write your own. - [Trust Engine](https://aceteam.ai/docs/safety-trust-engine): AceTeam's ensemble-of-judges approach to AI safety evaluation. Calibrated confidence scores from diverse judge models, validated by academic research. - [Hosted Instances](https://aceteam.ai/docs/instances-overview): Deploy AI agents as hosted containers with automatic safety enforcement. No infrastructure to manage: create an instance, pick a template, and your agent runs 24/7 with cost tracking and audit trails. - [Quickstart: Deploy Your First Instance](https://aceteam.ai/docs/instances-quickstart): Create a hosted AI agent instance in under a minute. Pick a template, set a safety policy, and your agent is live with full safety enforcement. ### Protocol - [Agentic Execution Protocol](https://aceteam.ai/docs/aep-overview): The protocol layer that makes multi-organization AI workflows accountable. Start here to understand why AEP exists and how to read the full specification. - [Agent Compute Protocol](https://aceteam.ai/docs/acp-overview): The protocol that lets AI agents dynamically request compute resources at runtime. ACP is the compute substrate counterpart to AEP's accountability layer. - [AEP Whitepaper](https://aceteam.ai/docs/aep-whitepaper): The complete Agentic Execution Protocol specification: cost accountability, provenance, and data governance for multi-organization AI workflows. ### API Reference - [Authentication](https://aceteam.ai/docs/authentication): Generate API keys and authenticate requests to the AceTeam API. - [SDK & API Quickstart](https://aceteam.ai/docs/sdks): Manage sandboxes, GPU compute, and ACET tokens from Python and TypeScript. - [Endpoints Overview](https://aceteam.ai/docs/endpoints-overview): Explore available API endpoint groups and request/response formats. - [Gateway API](https://aceteam.ai/docs/gateway): Route any OpenAI-compatible client through the AceTeam gateway for cost tracking, safety detection, model routing, and failover. ### FAQs - [General FAQs](https://aceteam.ai/docs/general-faqs): Common questions about pricing, models, privacy, and team management. - [Technical FAQs](https://aceteam.ai/docs/technical-faqs): Answers about rate limits, file sizes, WebSockets, and self-hosting. ### Company - [AceTeam Whitepaper](https://aceteam.ai/docs/accountable-compute-whitepaper): Sovereignty, cost transparency, and data governance in the agent economy. Why AI infrastructure needs an accountability layer, and how AceTeam built one. - [Strategic Overview](https://aceteam.ai/docs/strategic-overview): AceTeam: The Operating System for Sovereign AI --- # Documents ## Agent Builder Source: https://aceteam.ai/docs/agent-builder Configure AI agents with the Workbench: models, prompts, parameters, and tools. The Workbench is your primary workspace for creating, configuring, and testing AI agents. It provides a single interface where you select a model, write instructions, adjust behavior parameters, and iterate on your agent through live conversation before publishing it to your organization. ## Choosing a Model Every agent starts with a model selection. AceTeam supports multiple AI providers, and each model has different strengths: - **OpenAI GPT-4o** -- Strong general-purpose reasoning with multimodal capabilities (text, images, audio). Good default for most use cases. - **Anthropic Claude** -- Excels at nuanced writing, careful analysis, and following detailed instructions. Well suited for research and document-heavy tasks. - **Google Gemini** -- Handles long context windows effectively, useful when agents need to process large documents or extended conversation histories. - **Deepseek** -- Cost-effective option for high-volume tasks where budget matters. Performs well on coding and structured reasoning. You can change the model at any time without losing the rest of your configuration. This makes it easy to compare how different models handle the same prompt. ## Writing a System Prompt The system prompt defines your agent's identity, role, and behavior. It is the first message the model receives before any user input, and it shapes every response the agent produces. Effective system prompts typically include: - **Role definition** -- Who the agent is and what it does. For example: "You are a financial analyst specializing in SaaS metrics." - **Behavioral guidelines** -- How the agent should respond. Be specific about tone, format, and boundaries. - **Knowledge scope** -- What the agent should and should not attempt to answer. - **Output format** -- Whether the agent should use bullet points, tables, structured JSON, or conversational prose. Keep prompts direct and concrete. Vague instructions like "be helpful" produce generic behavior. Specific instructions like "always include the source document name when citing data" produce consistent, useful output. ## Tuning Parameters Below the system prompt, the Workbench exposes model parameters that control response behavior: - **Temperature** (0.0 to 1.0) -- Controls randomness. A value of 0 produces deterministic, focused responses. A value of 1 produces more varied and creative output. For factual tasks like data extraction, use 0 to 0.3. For brainstorming or creative writing, try 0.7 to 1.0. - **Max Tokens** -- The maximum length of each response. Set this based on how long you expect outputs to be. Short Q&A agents might need 256 tokens; agents that produce reports might need 4096 or more. - **Top-p** (nucleus sampling) -- An alternative to temperature that limits the model to the most probable tokens. A value of 0.9 means the model considers tokens that together account for 90% of the probability mass. In most cases, adjusting temperature alone is sufficient. ## Testing in the Chat Interface The right side of the Workbench provides a live chat panel. Type messages and see how your agent responds with the current configuration. This feedback loop lets you refine the system prompt and parameters iteratively: 1. Send a message that represents a typical user query. 2. Review the response for accuracy, tone, and format. 3. Adjust the system prompt or parameters. 4. Send the same message again to compare. You can clear the conversation at any time to start fresh, which is useful when testing how the agent handles an initial interaction without prior context. ## Saving and Versioning Every time you save an agent, AceTeam creates a version snapshot. Previous versions are preserved and accessible, so you can compare changes over time or roll back if a new prompt performs worse than the last one. Version history tracks the system prompt, model selection, and all parameters. ## Publishing to Your Organization Once you are satisfied with your agent's behavior, publish it to make it available to other members of your organization. Published agents appear in the agent directory where teammates can start conversations, embed agents into workflows, or fork them to create their own variations. Unpublished agents remain private to your account and do not appear to other organization members. --- ## Agent Tools & MCP Source: https://aceteam.ai/docs/agent-tools-mcp Extend agents with tools using the Model Context Protocol. By default, AI agents can only generate text. Tools give agents the ability to take actions -- query databases, call APIs, run workflows, and retrieve documents. AceTeam uses the Model Context Protocol (MCP) as the standard for connecting tools to agents. ## What Is MCP? The Model Context Protocol is an open standard that defines how AI models discover and invoke tools. Instead of hard-coding tool integrations into each agent, MCP provides a uniform interface: the agent sees a list of available tools with descriptions and parameters, decides when to call one, and receives structured results back. This means you can add new capabilities to an agent by connecting an MCP server, without modifying the agent's system prompt or code. ## Built-in Tool Categories AceTeam provides several categories of tools out of the box. When you enable a tool category on an agent, all tools in that category become available during conversations. ### Agent Tools Tools for interacting with other agents in your organization: - **list_agents** -- Returns a list of available agents with their names and descriptions. - **chat_with_agent** -- Sends a message to another agent and returns its response. This enables multi-agent collaboration where a primary agent delegates specialized tasks to others. ### Workflow Tools Tools for discovering and executing workflows: - **list_workflows** -- Returns available workflows and their descriptions. - **run_workflow** -- Triggers a workflow by name with input parameters and returns the result. This lets agents orchestrate multi-step automations mid-conversation. ### Document Tools Tools for accessing your organization's document store: - **list_documents** -- Returns documents matching a query or filter criteria. - **get_document** -- Retrieves the content of a specific document by ID. Useful for RAG (retrieval-augmented generation) scenarios where the agent needs to reference source material. ### Workbench Tools Utility tools for managing the agent's own environment, including conversation context management and configuration access. ## How Tools Appear to the Agent When a user sends a message, the agent receives the conversation history along with a list of available tools. Each tool includes a name, description, and a schema defining its parameters. The model decides whether a tool call is appropriate based on the user's request. For example, if a user asks "What workflows do we have for onboarding?", an agent with workflow tools enabled will: 1. Recognize that the question requires data it does not have in memory. 2. Call `list_workflows` with a relevant filter. 3. Receive the list of matching workflows. 4. Compose a natural language response summarizing the results. Tool calls happen transparently within the conversation. The user sees the final answer; the tool invocation details are available in the conversation log for debugging. ## Connecting External MCP Servers Beyond the built-in tools, you can connect any MCP-compatible server to extend an agent's capabilities. This is how you integrate AceTeam agents with your own systems. To connect an external MCP server: 1. Open the agent configuration in the Workbench. 2. Navigate to the Tools section. 3. Select **Add MCP Server** and provide the server URL. 4. AceTeam will discover the tools exposed by that server and display them in the tool list. 5. Enable the tools you want the agent to access. External MCP servers run on your infrastructure. AceTeam routes tool calls through the platform's backend, so your MCP server does not need to be publicly accessible -- it only needs to be reachable from the AceTeam Python backend or through the Sovereign Compute Fabric. ## Use AceTeam Tools from Any MCP Client You can connect AceTeam's MCP tools to Claude desktop, Codex, Claude Code, or any MCP-compatible client. This gives your AI assistant direct access to agent management, knowledge base search, workflow execution, and more. See [Connect AceTeam (MCP)](/docs/connect-mcp) for the connect URL and a recipe for each client. Claude Code speaks native remote MCP over HTTP with built-in OAuth: ```bash claude mcp add --transport http aceteam https://aceteam.ai/mcp ``` If you are an AI agent setting this up yourself, start at the [Agent Index](/docs/agent-index) (or [/llms.txt](/llms.txt)) for copy-pasteable install snippets and a machine-readable index of every doc. ## Best Practices - **Enable only what the agent needs.** Fewer tools means less ambiguity for the model when deciding which tool to call. - **Write clear tool descriptions.** The model relies on descriptions to understand when and how to use a tool. Vague descriptions lead to incorrect tool usage. - **Test tool interactions in the Workbench.** Use the chat panel to verify the agent calls the right tools with the right parameters before publishing. - **Monitor tool call logs.** Review conversation logs to identify cases where the agent misuses tools or fails to call them when it should. --- ## Onboarding External Harnesses Source: https://aceteam.ai/docs/external-harness-onboarding Point Codex, opencode, Cline, or a custom client at the AceTeam MCP and it self-hydrates identity, memory, and the playbook. # Onboarding External Harnesses Most coding agents carry context as a per-harness filesystem convention: Claude Code reads `MEMORY.md` and `CLAUDE.md`, Codex reads `AGENTS.md`, others have their own. AceTeam moves that context behind the one interface every harness already speaks, the MCP connection. Point any harness at the AceTeam MCP and it self-onboards: identity, curated memory, the how-to-work playbook, and the open work queue, all server-hosted. This page shows how to connect an external harness so a cold session hydrates on its own. ## How Hydration Works Two things carry the orientation, and they say the same thing: 1. **The MCP handshake.** The AceTeam MCP server sends a short instruction string in its `initialize` response. A client that surfaces server instructions to its model reads it at connect time. 2. **`AGENTS.md` at the repo root.** Harnesses that read a repo convention file instead of MCP server instructions (Codex is the primary case) get the same orientation from `AGENTS.md`, which reproduces the handshake string verbatim. Either path leads to the same first move: call the `bootstrap` tool once. It returns identity, the curated memory index version, the orchestrator playbook, the open worklist, and live guardrail baselines in a single call. If `bootstrap` is unavailable, call `memory_bootstrap` and read the `aceteam://playbook/orchestrator` and `aceteam://memory/index` resources. Then work the `automated`-labeled issues. ## Getting an API Key External harnesses authenticate with an `act_` API key. 1. Sign in to your AceTeam dashboard. 2. Go to Settings then API Keys. 3. Create a key and copy it once. Mint one key per harness, named for it (for example "Codex", "Claude Code", "ace-cli"), rather than sharing a single key across harnesses. The MCP server is mounted stateless, so the api key is the only durable per-harness attribution available server-side: `bootstrap.identity.api_key` reports the calling key's own name and explicit scopes back to the harness, and a shared key makes every harness using it indistinguishable there and in audit trails. Pass the key through an environment variable so it never lands in a committed file: ```bash export ACETEAM_API_KEY=act_REPLACE_WITH_YOUR_KEY ``` The endpoint is `https://aceteam.ai/mcp` (streamable HTTP), with the key sent as a Bearer token. ## Codex Add this to `~/.codex/config.toml`: ```toml [mcp_servers.aceteam] url = "https://aceteam.ai/mcp" bearer_token_env_var = "ACETEAM_API_KEY" ``` Codex reads the key from the `ACETEAM_API_KEY` environment variable, and reads `AGENTS.md` at the repo root on start. A copy-pasteable version of this config is committed at `docs/harness-onboarding/codex.config.example.toml`. ## opencode and Cline Clients that accept a streamable HTTP MCP server use the same endpoint with an `Authorization: Bearer` header: ```json { "mcpServers": { "aceteam": { "type": "http", "url": "https://aceteam.ai/mcp", "headers": { "Authorization": "Bearer act_REPLACE_WITH_YOUR_KEY" } } } } ``` Prefer sourcing the token from the environment where your client supports it, so the key stays out of the config file. ## Session Identity (Multi-Machine Addressing) The MCP server is mounted stateless: there is no connection-level session id, so nothing server-side can tell one Claude Code window on your laptop apart from another one on a Citadel node using the same API key. A session lease closes that gap (`session_register`, `session_list`, `whoami`'s `session` field), and it is addressed by a NAME your harness declares on every call through the `X-AceTeam-Session` header, plus the optional `X-AceTeam-Harness` and `X-AceTeam-Machine` headers. Claude Code, in `.mcp.json`: ```json { "mcpServers": { "aceteam": { "type": "http", "url": "https://aceteam.ai/mcp", "headers": { "Authorization": "Bearer act_REPLACE_WITH_YOUR_KEY", "X-AceTeam-Session": "${ACETEAM_SESSION:-}", "X-AceTeam-Harness": "claude-code" } } } } ``` launched with the session name in an environment variable: ```bash ACETEAM_SESSION=citadel claude ``` Codex, in `~/.codex/config.toml`, reading the header value from an environment variable rather than a literal: ```toml [mcp_servers.aceteam] url = "https://aceteam.ai/mcp" bearer_token_env_var = "ACETEAM_API_KEY" [mcp_servers.aceteam.env_http_headers] X-AceTeam-Session = "ACETEAM_SESSION" ``` launched the same way: ```bash ACETEAM_SESSION=ios codex ``` Either config is the recommended, no-code-change path: once it is in place, every AceTeam MCP call from that session carries the header, and the first call auto-registers the lease under that name. Codex sessions are also bound to their own `_meta.threadId`, present on every call with no configuration needed, so a Codex session that forgot to set the header can still call `session_register` explicitly and remain addressable. A name is scoped to your `(organization, user)` pair, so `citadel` on one machine and `ios` on another are siblings, never a collision, as long as they use different names. Registering the same name twice from a different declared machine refuses with the current holder's details; pass `takeover=true` to `session_register` to replace it (the common case is restarting a crashed session under the same name). Whether a header configured in a streamable HTTP client's `headers` block actually survives an OAuth-authenticated connection (as opposed to a direct `act_` Bearer key) has not been verified against a real interactive session as of this writing. If your session never shows up in `session_list` despite the header being set, call `session_register` explicitly instead; it works regardless of header delivery. ## Clients That Only Speak stdio For a client that cannot connect to a streamable HTTP server directly, bridge through `mcp-remote`: ```json { "mcpServers": { "aceteam": { "command": "npx", "args": ["-y", "mcp-remote@latest", "https://aceteam.ai/mcp"] } } } ``` The bridge handles auth via OAuth in the browser on the first tool call, so no inline key is required. ## Verifying a Cold Client Hydrates After connecting, confirm the conduit responds before relying on it. From the harness, or from any MCP client with the key: 1. Call `memory_bootstrap`. It returns a non-empty curated memory index (a markdown map grouped into Core, Infrastructure, Business, Products, Lessons, and Reference sections) with a provenance header naming the scope, version, and content hash. It never errors: if no curated head has been pushed for a scope, it serves a generated fallback synthesized from that org's memories. 2. Read the `aceteam://playbook/orchestrator` resource. It returns the orchestrator playbook markdown with a one-line provenance header. 3. Call `bootstrap` (if your client exposes it). It returns `identity`, `memory`, `playbook`, `worklist`, and `guardrails` in one object. A healthy cold client sees a non-empty memory index and a readable playbook on the first two calls without any hand-holding. If `memory_bootstrap` returns a generated fallback, the conduit is still working; a curated global head is maintained separately by the operator. --- ## Voice Agents Source: https://aceteam.ai/docs/voice-agents Enable real-time voice conversations and phone calls with your agents. AceTeam agents are not limited to text. With voice capabilities enabled, users can have real-time spoken conversations with agents directly in the browser or over the phone. This opens up use cases where typing is impractical or where a conversational interface is more natural. ## Browser-Based Voice Chat ### How It Works Voice conversations in the browser use WebRTC for low-latency audio streaming. The full pipeline looks like this: 1. **Audio capture** -- The browser captures microphone input using the Web Audio API. 2. **Speech-to-text** -- Audio is streamed to Deepgram, which transcribes speech to text in real time. 3. **AI processing** -- The transcribed text is sent to the agent's model, which generates a response. 4. **Text-to-speech** -- The response text is converted to audio using a TTS (text-to-speech) engine. 5. **Audio playback** -- The synthesized audio is streamed back to the browser and played through the speakers. This pipeline runs continuously during a voice session, creating a natural back-and-forth conversation. Latency is typically under two seconds from the end of a spoken phrase to the start of the agent's spoken reply. ### Enabling Voice on an Agent To add voice capabilities to an agent: 1. Open the agent in the Workbench. 2. Navigate to the Voice configuration section. 3. Enable **Voice Mode**. 4. Select a voice for the agent from the available options, or use the default. 5. Save and publish. Once enabled, a microphone button appears in the agent's chat interface. Users click it to start a voice session and click again to end it. Text chat remains available alongside voice. ### Tips for Voice Agents - **Keep system prompts concise in output.** Voice responses that run longer than 30 seconds feel slow. Instruct the agent to give brief, direct answers. - **Avoid markdown and formatting.** The TTS engine reads text literally. Bullet points, headers, and code blocks do not translate well to speech. - **Test with realistic audio.** Background noise, accents, and varied microphone quality affect transcription accuracy. Test in conditions that match your users' environment. ## Twilio Phone Integration Voice agents can be connected to phone numbers through Twilio, allowing users (or customers) to call a real phone number and speak with an AI agent. ### Setting Up Phone Access Twilio integration uses OAuth (no manual credentials required): 1. Navigate to the agent's Voice configuration. 2. Under **Phone Integration**, click **Connect Twilio**. 3. You will be redirected to Twilio to authorize AceTeam. Sign in and grant access. 4. Once connected, select the phone number you want to assign from your Twilio account. 5. Configure call handling behavior: - **Greeting message** -- What the agent says when it answers. - **Call timeout** -- Maximum call duration before automatic disconnection. - **Fallback behavior** -- What happens if the agent cannot process a request (transfer to human, leave a message, etc.). 6. Save the configuration. Incoming calls to the connected number will be routed to the agent. You can manage or disconnect your Twilio OAuth connection from **Organization Settings > Integrations**. Outbound calls are also supported. Agents within workflows can initiate phone calls as part of automated processes -- for example, calling a customer to confirm an appointment. ### Chat-Initiated Calls Agents can initiate phone calls directly from a chat conversation. When an agent determines that a phone call would be more appropriate (for example, to collect sensitive information or walk someone through a complex process), it can offer to call the user. The user provides a phone number in chat, and the agent places the call via Twilio while maintaining conversation context. ### Call Escalation When an AI agent reaches the limits of what it can handle, it can escalate the call to a human. Configure escalation rules in the agent's Voice settings: - **Transfer number** -- The phone number or queue to transfer to. - **Context handoff** -- The agent passes a summary of the conversation to the human agent, so the caller does not have to repeat themselves. - **Escalation triggers** -- Define conditions that trigger automatic escalation (e.g., sentiment detection, specific keywords, or explicit user request). ### Agent Call Control Agents can programmatically end calls when the conversation is complete. This is useful for automated scenarios like surveys, appointment confirmations, or notification calls where the agent should hang up after delivering its message and collecting a response. Call control actions (hang up, hold, transfer) are configured per agent in the Voice settings. ## Voice Cloning AceTeam supports custom voice creation so your agents can speak with a distinctive, branded voice instead of a generic TTS voice. ### Creating a Custom Voice Voice cloning is powered by ElevenLabs, delivering high-fidelity custom voices: 1. Go to the Voice configuration section. 2. Select **Custom Voice**. 3. Record or upload audio samples. For best results, provide at least 30 seconds of clear speech in a quiet environment. Multiple samples of varied sentences produce more natural output. 4. AceTeam processes the samples through ElevenLabs and generates a voice profile. 5. Assign the custom voice to any of your agents. Custom voices work with both browser-based voice chat and outbound Twilio calls. Voices are scoped to your organization: other organizations cannot access or use your voice profiles. ## Use Cases Voice agents are particularly effective in scenarios where hands-free interaction matters or where the target audience prefers speaking over typing: - **Customer service lines** -- Route inbound calls to an AI agent that handles common questions, escalating to a human when needed. - **Virtual assistants** -- Provide a voice interface for internal tools so employees can query systems while multitasking. - **Training simulations** -- Create realistic practice conversations for sales reps, support staff, or language learners. - **Appointment scheduling** -- Let customers call in, check availability, and book appointments without waiting for a human operator. - **Accessibility** -- Offer a voice-first experience for users who find text interfaces difficult to use. --- ## Ground Agents with Knowledge Collections (RAG) Source: https://aceteam.ai/docs/knowledge-collections Build a knowledge base with collections, upload documents and files, link a collection to an agent, and let it retrieve and cite that knowledge (RAG) during chat. Agents answer from their training data plus whatever context you give them. To ground an agent in your own material (contracts, product docs, meeting notes, research), you put that material in a knowledge collection and link the collection to the agent. From then on the agent retrieves relevant passages from the collection during chat and can cite where each passage came from. This is retrieval-augmented generation, or RAG. This page covers the full loop: create a collection, add documents and files, link it to an agent, and see how retrieval and source attribution work at chat time. Every step maps to an MCP tool, so you can do all of it from Claude Code, Cursor, or any MCP client (see [Agent Tools & MCP](/docs/agent-tools-mcp)), as well as from the web app. ## What a Collection Is A collection is a folder that holds searchable knowledge. Under the hood a collection, a folder, and a workspace are the same record, so the ID you get back from creating one works everywhere a folder ID is expected: uploads, sharing, and agent links. Because of that history, three tool names create the same thing: | Tool | Status | Notes | | --------------------- | ---------------- | ----------------------------------------------------------- | | `create_folder` | Canonical | Supports `name`, `description`, and `parent_id` for nesting | | `create_collection` | Deprecated alias | Kept working; takes `name` and `description` only | | `drive_create_folder` | Deprecated alias | Drive-framed name for the same operation | Prefer `create_folder`. The others still work so existing agents do not break. ## Create a Collection 1. Call `create_folder` with a `name`, and optionally a `description` and a `parent_id` to nest it under an existing folder. 2. Save the returned ID. This is the `collection_id` (also called `workspace_id` or `folder_id`) you will pass to upload and link tools. 3. Call `list_collections` at any time to list your organization's collections with their IDs and names. Collections are scoped to your organization. `list_collections` and `list_folders` both read the same set of workspace records for the calling org. ## Add Documents and Files to a Collection There are two distinct ways to put content into a collection, and the difference matters for what the agent can later do with it. ### Index Searchable Text: `upload_document` Use `upload_document` when the goal is for an agent to search and retrieve the content. | Parameter | Required | Default | Description | | ----------- | -------- | ------------ | ---------------------------------------------------- | | `title` | Yes | | Document title | | `content` | Yes | | Full text, or base64-encoded bytes for a PDF or DOCX | | `file_type` | No | `text/plain` | MIME type | | `folder_id` | No | | Collection (folder) to place the document in | The content is split into chunks and embedded so it becomes searchable. For a PDF or DOCX, pass the base64-encoded file bytes with the matching `file_type`; the text is extracted before chunking. Plain text formats are stored verbatim. Important: `upload_document` indexes the extracted text only. It does NOT keep the original file bytes. The document is fully searchable, but it is not a signable source, and it cannot feed a workflow step that needs to read the raw file. ### Store the Raw File: `upload_file` Use `upload_file` when you need the actual bytes preserved, for example to sign a PDF or DOCX later, or to feed a flow's file reader. `upload_file` stores the raw file in your AceTeam Drive (visible in `/browser`) and returns a scoped share link. Bytes come from exactly one of two sources: - `content`: base64-encoded bytes, for small blobs only (up to 50 KB of raw bytes) such as icons, tiny text, or thumbnails. - `node_id` plus `path`: the file is read from your own connected Citadel node over the mesh, which is the right path for a large local file. Pass `folder_id` to place the file in your collection. Other parameters include `mime_type` (inferred from the extension if omitted), `visibility` (`link` by default, or `org`, `users`, `private`), `expires_in_hours`, and `sha256` for an integrity check. Note that `upload_file` alone does not index the file for search. If you want a collection whose raw files are also searchable by the agent, you generally add both: `upload_file` for the bytes and `upload_document` for the searchable text. ### Choosing Between Them | | `upload_document` | `upload_file` | | ------------------------------------------ | ----------------- | ---------------------------------------- | | Stores raw bytes | No | Yes | | Indexes searchable text (chunk + embed) | Yes | No | | Feeds agent RAG retrieval | Yes | No (index it with `upload_document` too) | | Can back a signature or a flow file reader | No | Yes | | Destination | Knowledge base | AceTeam Drive (`/browser`) | A third tool, `drive_create_file`, writes into a connected external Google Drive account. Despite the similar name, it is unrelated to the two tools above. Pick by where the file needs to end up, not by the name. ## Link a Collection to an Agent Linking a collection to an agent is what scopes the agent's knowledge to that collection. 1. Call `link_collection_to_agent` with the `agent_id` and the `collection_id`. 2. Set `can_read` (default `true`) so the agent can read documents in the collection. Set `can_write` (default `false`) only if the agent should be able to add documents to it. At least one must be true. 3. To detach later, call `unlink_collection_from_agent` with the same pair. Both tools require organization membership and that you own both the agent and the collection. `link_collection_to_agent` is idempotent: re-linking the same pair updates the permissions in place rather than creating a duplicate. `unlink_collection_from_agent` is idempotent too: unlinking a collection that is not attached is a clean no-op. You can also wire knowledge at creation time by passing a `knowledge` item to `create_agent`, which performs the same link. To inspect an agent's current knowledge scope, call `get_agent`, which lists each linked collection with its name, ID, and `can_read` / `can_write` flags. ## How the Agent Retrieves and Cites Knowledge During Chat Once a collection is linked with read access, retrieval happens two ways. **Automatic injection every turn.** On each user message, the platform runs a bounded, top-k semantic retrieval against the agent's readable linked collections, scoped to that message, and injects the matches into the agent's context as supplementary information. This is the lowest-priority context source: it supports the answer without overriding the agent's instructions. It runs whether or not the model decides to search, so an attached collection shows up in context without the model having to ask for it. **Source attribution.** The injected block is labeled `Relevant knowledge:` and lists each retrieved passage numbered and prefixed with the document it came from (`From :`). That is what lets the agent name the source of a fact in its reply, so you can trace an answer back to the document behind it. To reinforce this, you can instruct the agent in its system prompt to always name the source document when it uses retrieved knowledge (see [Agent Builder](/docs/agent-builder)). **On-demand deep search.** The agent also has the `search_knowledge_base` tool available for a deeper lookup when the automatically injected passages are not enough. The two paths are independent: automatic injection keeps recent, relevant context present every turn, while the tool lets the agent search on demand. ## Search the Knowledge Base Directly You (or the agent) can query the knowledge base at any time with `search_knowledge_base`. | Parameter | Required | Description | | --------------- | -------- | ---------------------------------- | | `query` | Yes | Search query text | | `collection_id` | No | Scope the search to one collection | Search is hybrid: it combines vector (semantic) similarity over embedded chunks with word-level text matching as a fallback. Results are deduplicated, and vector matches are prioritized. Passing a `collection_id` restricts the search to that collection and its nested subfolders, which is the same scoping that a linked collection applies to an agent's searches. ## Manage Documents and Collections - `list_documents` (alias for `file_list`) lists documents in your organization, optionally filtered by `folder_id`. - `get_document` (alias for `file_get`) retrieves a document's details and content preview. - `delete_document` (alias for `file_delete`) soft-deletes a document. It uses two-step confirmation: the first call (`confirm=false`, the default) returns a preview and changes nothing; call again with `confirm=true` to delete. Deletions are soft by default, in line with the platform's data-ownership model: the record is marked deleted rather than dropped. ## Typical Flow 1. `create_folder` to make the collection. 2. `upload_document` for each piece of text you want searchable (add `upload_file` as well if you need the raw bytes). 3. `link_collection_to_agent` with `can_read=true`. 4. Chat with the agent. It retrieves relevant passages automatically and can cite the source document, and it can call `search_knowledge_base` for deeper lookups. --- ## Deploying Agents to Slack, WhatsApp, and Telegram Source: https://aceteam.ai/docs/channels-messaging Connect an agent to Slack, WhatsApp, Telegram, Discord, email, and SMS or phone. Set up a WhatsApp bot, a Telegram bot, and deploy an agent to messaging channels so it replies. AceTeam agents are not confined to the web chat. You can connect an agent to the messaging channels your team and customers already use: Slack, WhatsApp, Telegram, Discord, email, and SMS or phone. Each channel is connected once at the organization level, after which your agents can read and send messages there. Some channels also deliver inbound messages straight to an agent so it replies automatically. There are two things worth separating up front: - **Connecting a channel** stores a credential (an OAuth token, a bot token, or a bridge URL and key) for your organization. This is what makes the channel's tools work. - **An agent responding** happens in one of two ways. On channels with an inbound hook (Slack, Telegram, phone), a message or call is routed to a specific agent that replies on its own. On the other channels, an agent participates by being given the channel's read and send tools and then invoked from chat, a scheduled job, or a Flow. All connect, send, and read operations are exposed as MCP tools, so an agent (or you, from any MCP client) can run them. Several channels also have a Connected Accounts panel in the web app under Integrations. ## Slack Slack connects through OAuth and stores one bot token for your whole organization. Once connected, mentioning the bot in a channel triggers an agent reply. 1. Call `slack_connect` (or click Connect Slack under Integrations). It returns a Slack OAuth URL. 2. Open the URL and authorize AceTeam for your workspace. You are redirected back when Slack approves. 3. Verify with `slack_status`, which reports whether Slack is connected for the current organization. 4. Invite the bot to the channels you want it in, and give it read and post permissions. Which agent replies: when the bot is mentioned, AceTeam resolves the agent attached to that Slack workspace. If a specific agent was bound when the workspace was connected, that agent responds; otherwise it falls back to your organization's most recently created agent. If the workspace is connected but no agent exists yet, the bot posts a short notice asking you to connect an agent rather than staying silent. Agent tools once connected: `slack_send_message` (post to a channel or thread, attributed to the triggering user), `slack_read_messages`, `slack_list_channels`, and `slack_status`. Disconnect with `slack_disconnect`, which clears the stored bot token for the whole organization and asks Slack to revoke it. ## Telegram Telegram offers two paths. The shared bot needs no setup; bringing your own bot gives you a branded bot. Both end with people messaging the bot and an agent replying. ### Shared bot (no setup) 1. Call `telegram_link`. It mints a one-time deep link (valid about 10 minutes) scoped to your user and organization. 2. Open the link in Telegram and tap Start. That chat is now bound to your workspace, and your organization's default (Ace) agent replies to messages in it. 3. Inside the chat, send `/switch` to change workspace or `/stop` to unlink. List your linked chats with `telegram_list_links` and remove one with `telegram_revoke_link`. ### Your own bot 1. Create a bot in Telegram with @BotFather: message it, send `/newbot`, choose a display name, then a username that ends in `bot` (5 to 32 characters, letters, digits, and underscores only, no hyphens). Copy the token it returns. 2. Call `telegram_connect` with that `bot_token`. AceTeam validates it against Telegram's `getMe` and stores it encrypted, so a bad token fails fast. 3. Call `telegram_set_webhook` to turn on inbound. This tells Telegram to push every incoming message to AceTeam and wires the bot to your organization's Ace agent, so replies are automatic. 4. Message the bot. The full order is connect bot, set webhook, then DM the bot. Note: Telegram makes webhooks and polling mutually exclusive. Once you set the webhook, the `telegram_read_messages` and `telegram_list_chats` pull tools stop returning new updates because inbound is now event driven. Use `telegram_disable_webhook` to return to pull mode. Check connection and delivery health with `telegram_status` and `telegram_health`. ## WhatsApp The WhatsApp MCP tools drive a self-hosted, headless bridge (Baileys) that logs in by scanning a QR code with your phone. This is a different system from the WhatsApp Business API channel (Twilio or Meta) that also appears in the Integrations UI; the tools below are the bridge. There are two ways to attach a bridge, and provisioning is a separate step from connecting. ### Connect a bridge you already run 1. Stand up a WhatsApp bridge that AceTeam can reach. A public address or a Headscale mesh host works directly; other private ranges require the operator to allow private networks. 2. Call `whatsapp_connect` with the bridge `api_url` and `api_key`. By default the credential is scoped to you alone; pass `org_level=true` to store one shared bridge any member can use. 3. Open the bridge's `/qr` endpoint and scan the code with WhatsApp to log the number in. 4. Confirm with `whatsapp_status` (returns the masked bridge URL) and `whatsapp_health` (live login state and which number is logged in). ### Provision a bridge on your own node If you run a Citadel fabric node, you can have AceTeam deploy the bridge for you. Provisioning both deploys the bridge and stores its credential, so it subsumes the connect step above. 1. Call `whatsapp_provision` with the `node_id` of the Citadel node to host it (optionally a `tenant` name and outbound `proxy`). This requires an API key carrying the `node:whatsapp` scope, which is narrower than shell access. 2. The first call previews and changes nothing. Call again with `confirm=true` to dispatch. Provisioning is fail-closed: if the target node's worker is not reachable, it refuses rather than deploy to another node. 3. The node deploys the Baileys bridge container, mints a per-tenant API key, and returns the bridge URL, key, and a pairing QR. AceTeam stores that credential as your organization's WhatsApp credential automatically. 4. Scan the returned QR with WhatsApp to log the number in. Agent tools once connected: `whatsapp_send_message` (to a JID, a phone number, or an exact group name), `whatsapp_read_messages`, `whatsapp_list_contacts`, `whatsapp_list_groups`, and `whatsapp_download_media`. On the bridge, an agent participates by holding these tools and being invoked; there is no automatic inbound reply hook for the bridge. Disconnect with `whatsapp_disconnect`, which logs the session out and clears the credential. ## Discord Discord connects with a bot token from the Discord Developer Portal. 1. In the Developer Portal, create an application, add a Bot, and reset or copy its token. 2. Call `discord_connect` with that `bot_token`. AceTeam validates it against Discord's `GET /users/@me` and stores it encrypted. 3. Invite the bot to your server with the `bot` scope and Read and Send Messages permissions. 4. Confirm with `discord_status` and `discord_health`. Agent tools once connected: `discord_send_message`, `discord_read_messages`, `discord_list_channels`, `discord_status`, and `discord_health`. An agent participates by holding these tools and being invoked. Automatic inbound replies on Discord are handled through a separate server-install flow and are limited to slash-command style interactions in a bound server, so a plain channel message does not trigger an automatic agent reply through `discord_connect` alone. ## Email Connecting an email account lets an agent read, search, draft, and send mail on your behalf. It is not an auto-responder; the agent acts when you invoke it. 1. Call `email_connect`. It returns a URL to the Integrations page (Emails tab) where you authorize an account. Multiple accounts (personal, work) can be connected. 2. Complete the provider authorization in the browser. 3. Verify with `email_list_accounts`, which lists connected accounts, their labels, and which is the default. Agent tools once connected: `email_search`, `email_read`, and `email_send` and `email_reply` for outbound. Note that `email_draft` does not create a provider draft; it queues the message at `/approvals` for a human to approve before it is sent, which is the safe default for agent-composed mail. ## SMS and phone Voice and SMS run on a connected phone provider. Buying numbers and sending SMS are organization-admin actions, and an agent answers inbound calls only after you explicitly attach it to an extension. For deeper voice configuration, see [Voice Agents](/docs/voice-agents). 1. Connect a provider with `phone_provider_connect`. For Twilio, pass `account_sid` and `auth_token`; for VoIP.ms, pass `api_user` and `api_pass`. Credentials are verified with a free read-only call before being stored encrypted. The web Twilio Connect (OAuth) flow is the recommended path for production. 2. Find a number with `phone_number_search`, then buy it with `phone_number_buy`. Because this spends money, the first call previews and you must call again with `confirm=true`. 3. Attach an agent to answer calls with `phone_attach_agent`, passing the `agent_id`, the number (`e164`), and an optional `extension` (auto-allocated from 101 if omitted). Callers dial your number, hear the IVR, enter the extension, and reach that agent. An organization that owns no number can attach to the shared AceTeam platform number (`+15484909240`) to receive calls with no setup. Detach later with `phone_detach_agent`. 4. Send outbound texts with `phone_send_sms`, which auto-resolves the sender to your organization's active number (or the shared platform number for a confirmed human session) and can be routed through the approval gate. ## Disconnecting a channel Each channel has its own teardown: `slack_disconnect`, `whatsapp_disconnect`, `telegram_revoke_link` (shared bot) or `telegram_disable_webhook` (your own bot), `phone_detach_agent`, and `google_disconnect` for Google-backed email. Disconnecting clears the stored credential so no agent can read or post through that channel until it is connected again. --- ## Giving Agents Persistent Memory Source: https://aceteam.ai/docs/agent-memory How agent memory works: store and recall facts across sessions, org-scoped memory files, automatic recall, pinning, and inspecting what an agent remembers. # Giving Agents Persistent Memory Agents remember nothing between conversations unless you give them memory. AceTeam agents have two distinct memory systems, and they behave differently. One is injected into the agent automatically on every turn. The other is a store of memory files the agent looks up on demand. This page explains how each one stores and recalls facts across sessions, how memory is scoped, and how you inspect or manage what an agent remembers. ## Two memory systems The word "memory" covers two unrelated stores. Knowing which one you are working with is the whole game. | System | Written by | Read into a turn | Best for | | ---------------- | ---------------------------------------- | ------------------------------------------------- | -------------------------------------------------- | | Injected hot-set | `agent_memory_set`, automatic extraction | Automatically, every turn | Short facts that should shape every reply | | Memory files | `memory_write` | On demand, via `memory_search` then `memory_read` | Reference material an agent looks up when relevant | The injected hot-set is a set of one-line facts loaded in front of the model before it answers. Memory files are markdown documents the agent has to go read. Writing to one does not populate the other. ## The injected hot-set: automatic recall Every time an agent takes a turn, the platform loads a bounded set of remembered facts directly into the agent's context. Nothing has to call a tool for this to happen. This is what makes an agent feel like it remembers you. ### How facts get stored There are two ways a fact enters this store. - **Automatic extraction.** After a conversation turn, the platform runs a lightweight model over the exchange, pulls out high-confidence facts, preferences, and patterns, and saves them. This is throttled and best-effort, so it never blocks or fails the user's turn. You do not have to do anything to get it. - **Explicit writes.** Call `agent_memory_set` with an `agent_id` and a plain-language `fact` (for example, "Prefers concise answers with no filler") to write a fact directly. You need write permission on the agent, since this changes what that agent knows. An optional `importance` value from 0.0 to 1.0 acts as a tiebreaker in ranking. The call returns the new memory's id so you can pin it. ### How recall works At the start of each turn the platform selects a "hot-set" that fits a fixed byte budget of about 4KB, so the loaded set stays small no matter how many memories an agent has accumulated. Selection runs in this order: 1. **Pinned memories first.** A pinned fact is always loaded, even if the pinned set alone exceeds the budget. A pin is an override. 2. **A slice reserved for the current topic.** About a third of the budget is held for facts whose keywords match the current message, so a fact that is relevant right now loads even if it is old or rarely used. 3. **The rest by recency and frequency.** Remaining space is filled by a score that combines how recently and how often each fact was recalled. Recency decays with a roughly 30-day half-life. When a fact is loaded into a turn, its recall counters are nudged so the hot-set drifts toward what is actually used. Facts that do not make the cut are never deleted. They stay available and can be surfaced again later. ### Pinning a fact Use `memory_pin` with the memory's id to keep a durable fact (a standing preference, a project invariant) loaded regardless of budget or ranking. Pass `pinned: false` to unpin, which returns the fact to normal recency-and-frequency ranking. The id is a recalled-memory row id, not a memory-file slug. You can only pin a fact owned by your own organization and user; anything else reports as not found and changes nothing. ### Scope Hot-set facts are scoped to a specific agent and user pair, within your organization. A fact one user's conversations produced is not injected for a different user, and facts written for one agent are not visible when a different agent runs. Note also that a fact written here with `agent_memory_set` is not mirrored into the knowledge base, so it will not appear in `memory_search` or `search_knowledge_base`. Automatically extracted facts are mirrored, so two facts can sit side by side in this store with different discoverability. ## Memory files: org-scoped, looked up on demand The second system is a set of markdown memory files stored on your organization's primary Citadel node, under a memory sandbox in the node's workspace. These are shared across the organization and organized by scope rather than tied to one user. They are not injected automatically. An agent only sees a memory file when it (or you) explicitly searches for it and reads it. ### Scopes Every memory file lives in one scope: - `global` (the default) for organization-wide facts. - A project name, such as `aceteam`, for facts about one project. - `agent:{id}` for facts tied to a specific agent. ### The file tools | Tool | What it does | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `memory_write` | Create or update a memory file. Idempotent by slug and content: writing the same content again is skipped. Frontmatter is generated for you. Secrets in the content are scrubbed before anything is stored, and the write fails closed rather than storing unscrubbed secrets. | | `memory_read` | Read one memory file by its slug. With no scope, it searches global first, then projects, then agents. | | `memory_search` | Find memories by query. It runs a semantic (vector) search over your organization's memory documents first, then falls back to keyword and node search. Returns ranked snippet previews. | | `memory_list` | List memory files as a table, optionally filtered by scope. | | `memory_delete` | Soft-delete a memory. The file is overwritten with a tombstone rather than physically removed, preserving history. It uses two-step confirmation: the first call previews, and a second call with `confirm: true` performs the delete. | ### Durability How a write survives depends on your organization's storage policy. For organizations that mirror memory centrally, `memory_write` writes the central copy first and then syncs to the node. If the node is briefly unreachable, the write still succeeds and syncs to the node automatically on a later memory operation. Organizations set to keep memory strictly on their own node do not mirror centrally, so a node failure fails the write. ## Bootstrapping at session start Call `memory_bootstrap` at the start of a session to load your organization's curated memory index. It returns a single map of who the user is, active work, and standing rules, served as a versioned artifact so any client can start cold, not just one machine. Pass a `scope` (`global` by default, a project name, or `agent:{id}`) and a `format` of `markdown` or `json`. If no curated index has been pushed for a scope, it serves a generated fallback synthesized from your memories rather than returning an error. Prefer it over guessing, then use `memory_search` and `memory_read` for detail. ## Inspecting and managing memory You have a few ways to see what an agent remembers. - **Preview the hot-set.** `agent_preview_context` shows the exact context a real turn would inject, including the memory hot-set, without running the model or spending anything. It is strictly read-only: previewing never mutates recall counters or perturbs ranking. - **Browse memory files.** Use `memory_list` and `memory_read` to see and read the file store. - **The Agent memory browser.** Organization admins can open Settings, then Organization, then Memory to browse stored memory records. You can filter by namespace, kind, and provenance, and free-text search over labels. A detail view shows each record's provenance (who created and updated it, and its source) and usage counters. From here an admin can soft-delete a record, which tombstones it while preserving history. Personal, user-scoped records stay private to their owner and are not shown to admins. ## What is authoritative, and what is not Treat the ranking counters as hints, not an audit trail. The recall count and last-recalled timestamp on a hot-set fact are best-effort signals that steer which facts load. They are debounced and updated without ever failing a turn, so they are not an exact record of how many times a fact was used. They exist to make recall better over time, not to account for it precisely. When you need the real content of a memory, read the record itself rather than trusting a counter. --- ## Generate Images and Video on Your Own Fabric Source: https://aceteam.ai/docs/image-video-generation Render images and short video clips from a text prompt on your own Citadel node, with an optional bring-your-own-key cloud fallback for images. `generate_image` and `generate_video` render media from a text prompt, sovereign first: by default the prompt is dispatched to your organization's own Citadel node rather than a cloud image or video API. Both tools share the same dispatch core, the same `MEDIA_GENERATE` fabric job, and the same output path into your AceTeam Drive, so a result from either tool composes with `page_attach_file`, an email attachment, or any other tool that takes a Document ID. ## Image Generation `generate_image` dispatches your prompt to your organization's own Citadel node over the fabric, where a `diffusers` sidecar serves Stable Diffusion, SDXL, or Flux family models. Nothing is sent to a cloud image API by default. | Parameter | Required | Default | Description | | ------------------------ | -------- | ------------------ | -------------------------------------------------------------- | | `prompt` | Yes | | The text prompt to render | | `model` | No | node default | Model override, e.g. `stabilityai/stable-diffusion-3.5-medium` | | `negative_prompt` | No | | Optional negative prompt (fabric only) | | `width` / `height` | No | 512 / 512 | Output dimensions in pixels, 64 to 2048 (fabric only) | | `steps` | No | 4 | Denoising steps, 1 to 150 (fabric only) | | `seed` | No | | Seed for reproducible output (fabric only) | | `node_id` | No | any ready node | Citadel node to generate on | | `filename` / `folder_id` | No | `image.png` / Home | Where the result lands in your Drive | | `provider` | No | `auto` | `auto`, `fabric`, or `gemini` | `provider="auto"` is sovereign first with a cloud fallback: only when no image-generation-ready node is connected does it fall back to Gemini 2.5 Flash Image, and only if your organization has its own Google API key configured in the Credentials Vault. There is no platform-wide key fallback: a missing key fails with an actionable message to connect one. Pass `provider="fabric"` to disable the cloud fallback entirely, or `provider="gemini"` to skip the fabric probe and go straight to your own Google key. Sovereign generation is metered at zero cost, since it runs on hardware you already own and is never routed through a cloud cost tracker. The Gemini path is also zero platform cost: your own Google key pays Google directly. ## Video Generation `generate_video` mirrors the image tool's dispatch pattern exactly, targeting the node's text-to-video model (Wan2.1-T2V-1.3B by default, catalog name `wan2.1-t2v`) on the same `diffusers` sidecar and the same `MEDIA_GENERATE` job path. | Parameter | Required | Default | Description | | ------------------------ | -------- | ------------------ | -------------------------------------------------------------------- | | `prompt` | Yes | | The text prompt describing the video | | `model` | No | node default | Model override, e.g. an HF repo id | | `width` / `height` | No | 832 / 480 | Output dimensions in pixels, 64 to 1280 | | `num_frames` | No | 81 | Frames to render, 1 to 161 (roughly 5 seconds at the default fps) | | `fps` | No | 16 | Frame rate for the exported MP4, 1 to 60 | | `steps` | No | 50 | Denoising steps, 1 to 150 (Wan2.1 is not a distilled few-step model) | | `seed` | No | | Seed for reproducible output | | `node_id` | No | any ready node | Citadel node to generate on | | `filename` / `folder_id` | No | `video.mp4` / Home | Where the result lands in your Drive | | `provider` | No | `auto` | `auto` or `fabric`, currently identical | There is no cloud text-to-video provider today, so `provider="auto"` and `provider="fabric"` behave the same: sovereign fabric only. The `provider` parameter exists so a future cloud fallback can slot in later without a rename, the same way Gemini was added for images. Video generation is currently a capability in progress on the node side. The sidecar's `/generate/video` endpoint exists, but the Citadel job handler that bridges a `MEDIA_GENERATE` job to it has not shipped as a separate follow-up yet. Until a node advertises a text-to-video-ready `diffusers` service, `generate_video` fails soft with a clear, retryable message rather than hanging or surfacing a raw dispatch error, and it names the gap so you know to connect a node with that service or try again later. When an explicit `node_id` is given and that specific node cannot serve the request, the error is about that node rather than the whole organization. Video generation is metered at zero cost as well, for the same reason as image generation: it runs on hardware you already own. ## Where the Result Lands Both tools store their output in your AceTeam Drive through the same path `upload_file` uses, so the rendered file shows up in `/browser` like any other Drive file. The tool result includes the share URL and the Document ID, which you can pass directly to `page_attach_file` to attach it to a hosted page, or to an email tool's attachment list. --- ## Author and Install Reusable Agent Skills Source: https://aceteam.ai/docs/skills Package instructions and an optional bundled Flow into a versioned, installable skill, then pin or track its updates per agent. A skill packages a set of instructions, and optionally a bundled Flow, into a named, versioned capability you can install on an agent. Skills are versioned the same way a Flow is: authoring a new skill version is how you adopt a change, whether that is edited instructions or a newer Flow graph. **This capability is authored but not yet enabled on this deployment.** The underlying tables (`skills`, `skill_versions`, `resource_skills`) are proposed in a migration that is awaiting founder schema approval and has not been applied yet. Every skill tool below degrades to a clean "not connectable" refusal until that migration is applied; nothing raises a raw error. Treat this page as describing the intended shape of the feature rather than something live today. Also explicitly out of scope for this slice, regardless of the schema: there is no `/skills` library page in the web app, no always-in-context injection of installed skills during a live agent turn, and no native `use_skill` runtime tool that would actually execute a skill's bundled Flow or surface its instructions mid-turn. `resolve_agent_skills` (below) is the resolution read a future runtime tool would call; it is exposed directly over MCP now so the underlying primitive is testable end to end ahead of that later wiring. ## Authoring a Skill: `skill_create` | Parameter | Required | Default | Description | | ----------------- | -------- | ------------------ | ----------------------------------------------------------------------------- | | `title` | Yes | | Skill title | | `description` | Yes | | One-line description shown wherever the skill is listed | | `instructions` | Yes | | The markdown instruction body | | `name` | No | derived from title | Invoke slug, lowercase letters, digits, and underscores, unique per org | | `trigger_hint` | No | | A short "use when..." hint | | `flow_id` | No | | An existing Flow in your org to bundle | | `flow_version_id` | No | current version | Pin to a specific Flow version instead of the current one, requires `flow_id` | | `entrypoint` | No | see below | `instructions` or `flow` | | `visibility` | No | `org` | `private`, `org`, or `link` | A skill is invoked by its `name`. Bundling a `flow_id` makes the Flow the skill's default entrypoint (`flow`) unless you pass `entrypoint="instructions"` explicitly, which keeps the Flow wired but runs the instructions first. A pure-instruction skill with no `flow_id` is always instruction-entrypoint; passing `entrypoint="flow"` without a `flow_id` is rejected. Free-tier organizations may author up to about three live skills, and any skill they author is forced to `visibility="org"` regardless of what is requested. Pro organizations and above are unlimited and may also use `visibility="link"`, the later share and marketplace surface. This tier gate applies only to authoring; installing or resolving a skill is never gated by tier. ## Installing a Skill on an Agent `install_skill_on_agent(agent_id, skill_id, pin_version_id?)` attaches a skill to one of your agents. A same-organization install defaults to track-latest: the agent always resolves the skill's current published version at read time, until you pin it explicitly with `skill_pin_version`. A skill owned by a different organization, such as a marketplace or cross-org skill, always pins at install time to a specific version, resolved automatically to the skill's current version unless you name one with `pin_version_id`. A private skill owned by another organization cannot be installed cross-org at all. `uninstall_skill_from_agent(agent_id, skill_id)` is the exact inverse and is idempotent: uninstalling a skill that is not currently installed is a clean no-op. ## Pinning and Unpinning `skill_pin_version(agent_id, skill_id, version_id?)` is the one-click pin for a same-org install that defaulted to track-latest. Omit `version_id` to pin to the skill's current version, freezing it in place, or pass one to pin to an older version instead. `skill_unpin_version(agent_id, skill_id)` reverts an installed skill to track-latest. It is refused for a cross-org or marketplace install, which always stays pinned by design; use `skill_pin_version` to move that pin to a different version instead. ## Reading a Skill `skill_get(skill_id)` returns a skill's current version: title, description, trigger hint, visibility, revision, entrypoint, and the full instruction body. ## Resolving What's Installed on an Agent `resolve_agent_skills(agent_id)` resolves every skill installed on an agent to its concrete version, the same resolution a future runtime tool would perform: a pinned install resolves to its pinned version, and a track-latest install resolves through the skill's current revision pointer. The result is a table of each installed skill's name, title, trigger hint, entrypoint, resolved version, and whether it bundles a Flow. --- ## Steer a Browser on a Citadel Node with Human Handoff Source: https://aceteam.ai/docs/cobrowse Drive a headed browser session on your own hardware from an agent, hand control to a human for login or 2FA, and audit every scripted action. Co-browse lets an agent steer a real, headed browser session running on one of your Citadel nodes, with a human able to take over at any point to log in, complete two-factor authentication, or otherwise interact directly. There are two distinct capabilities under this name, and most tools support both. ## Two Session Models **Single-session (legacy).** Every co-browse tool below also works with no `session_id` argument at all, in which case it drives the node's original single-session browser (the `COBROWSE` job). A human takes over by watching and interacting with the node's existing VNC or desktop stream. **Multi-session (current).** `cobrowse_start` launches a new, isolated session (the `COBROWSE_SESSION` job) and returns a node-issued `session_id`. Pass that `session_id` to the other tools to target that specific session rather than the legacy single browser. Each call to `cobrowse_start` launches a new session; it is not idempotent, so stop a session with `cobrowse_stop` when you are done with it. A multi-session launch can also unlock an encrypted, persistent browser profile by passing a `pin`, so logins survive across sessions instead of starting logged out every time. The PIN is forwarded to the node for that one call only. It is never logged, never stored in the session or audit record, and never included in any error message. Omitting `pin` launches a throwaway, logged-out session. Some scripted actions, click, type, and extract, exist only on the multi-session path; the original single-session job never grew them. ## Tools | Tool | Session models | Description | | --------------------- | ------------------ | -------------------------------------------------------- | | `cobrowse_start` | multi-session only | Launch a new session, returns its `session_id` | | `cobrowse_status` | both | Query session state, current URL, and who is driving | | `cobrowse_stop` | multi-session only | Tear down a session by its `session_id` | | `cobrowse_navigate` | both | Navigate to a URL | | `cobrowse_handoff` | both | Pause AI control, hand it to a human | | `cobrowse_resume` | both | Return control to the agent | | `cobrowse_screenshot` | both | Capture the current viewport | | `cobrowse_click` | multi-session only | Click by CSS selector or by `x`/`y` coordinates | | `cobrowse_type` | multi-session only | Insert text into the currently focused element | | `cobrowse_extract` | multi-session only | Read an element's text and, optionally, named attributes | For every tool that accepts `session_id`, omit it to act on the legacy single-session browser instead. ## Human Handoff and Driver Arbitration Any scripted action against a session refuses with a "handed off" error while a human is driving it, either because a viewer is currently watching the session live, or because `cobrowse_handoff` was called explicitly for it. An explicit handoff is sticky: it survives a transient viewer disconnect, such as a brief network blip mid-login, and clears only when `cobrowse_resume` is called for the same session. A viewer simply attaching to watch also blocks scripted actions for as long as it is attached, but that block clears on its own the moment the viewer detaches, with no explicit handoff call needed. `cobrowse_status` reports the current driver, `ai` or `human`, along with the session's current URL and a URL-based heuristic for whether the page looks signed in, so an agent can decide whether to hand off for a login before attempting to steer. ## Write Consent and Session Ownership Writes, `cobrowse_click`, `cobrowse_type`, and `cobrowse_navigate` (both the session and legacy forms), require a human-granted, per-session, time-boxed write consent grant before they may dispatch. Absent an active grant, the tool returns an approval envelope instead of acting; once a human approves it, retry the call. Reads, `cobrowse_screenshot`, `cobrowse_extract`, and `cobrowse_status`, never require this grant, though `cobrowse_screenshot` and `cobrowse_extract` are still refused while a human is actively driving, since a round trip into a page a human is using is not treated as a passive read. Control-plane calls, `cobrowse_start`, `cobrowse_stop`, `cobrowse_handoff`, and `cobrowse_resume`, are never gated by a write grant either. Every tool that takes a `session_id` also requires that you are the same user who started that session with `cobrowse_start`. Another member of your organization cannot drive, stop, or inspect a session you started, even though your organization as a whole owns or shares the node; they get a clear refusal naming the mismatch instead. This ownership check runs before the write consent gate, so a non-owner can never ride a grant a human approved for the session's actual owner. ## Audit Trail Every action against a session, read or write, successful or refused, is recorded to a computer-use audit trail: who acted, what was intended, and what actually happened. `cobrowse_screenshot` additionally attaches the captured image to its audit row. The audit write is best effort and never blocks or unwinds an already-dispatched action if it fails, but a failure to write it is always logged, never silently swallowed. --- ## Authentication Source: https://aceteam.ai/docs/authentication Generate API keys and authenticate requests to the AceTeam API. # Authentication All requests to the AceTeam API require authentication via an API key. This guide covers how to generate keys, authenticate requests, and follow security best practices. ## Generating an API Key 1. Sign in to your AceTeam dashboard. 2. Navigate to **Settings > API Keys**. 3. Click **Create New Key**. 4. Give your key a descriptive name (for example, "Production Backend" or "CI Pipeline"). 5. Copy the key immediately. For security reasons, the full key is only displayed once. ## Key Format API keys follow the format: ``` act_<64-character-hex-string> ``` For example: ``` act_a1b2c3d4e5f6... ``` All keys are prefixed with `act_` so they can be easily identified in configuration files and secret scanners. Keys are encrypted at rest and scoped to the organization that created them. A key created by a member of Organization A cannot access resources belonging to Organization B. ## Authenticating Requests Include your API key in the `Authorization` header of every request using the Bearer scheme: ```bash curl -X GET https://aceteam.ai/api/v1/agents \ -H "Authorization: Bearer act_your_key_here" \ -H "Content-Type: application/json" ``` If the key is missing, expired, or invalid, the API returns a `401 Unauthorized` response. ## Rate Limits API requests are subject to rate limits that vary by plan. When you exceed the limit, the API returns a `429 Too Many Requests` response. Rate limit details are included in response headers: | Header | Description | | ----------------------- | ---------------------------------------- | | `X-RateLimit-Limit` | Maximum requests allowed per window | | `X-RateLimit-Remaining` | Requests remaining in the current window | | `X-RateLimit-Reset` | Unix timestamp when the window resets | Implement exponential backoff in your client to handle rate limit responses gracefully. ## Key Rotation Rotating keys regularly reduces the risk of compromised credentials. Follow this process to rotate without downtime: 1. **Create a new key** in Settings > API Keys. 2. **Update your clients** and services to use the new key. 3. **Verify** that all systems are functioning correctly with the new key. 4. **Revoke the old key** by clicking the delete icon next to it in the dashboard. There is no limit on the number of active keys per organization, so you can run both keys simultaneously during the transition period. ## Security Best Practices - **Never commit keys to version control.** Use a `.gitignore` rule to exclude files containing secrets. If a key is accidentally committed, revoke it immediately and generate a new one. - **Use environment variables.** Store keys in environment variables or a secrets manager rather than hardcoding them in source files. - **Restrict access.** Only share keys with team members and systems that need them. Use separate keys for development, staging, and production environments. - **Monitor usage.** Review the API Keys page periodically to identify unused keys and revoke them. - **Use secret scanning.** Enable secret scanning in your repository hosting provider to detect accidentally committed keys. The `act_` prefix is compatible with most scanning tools. If you believe a key has been compromised, revoke it immediately from the dashboard. Revoking a key takes effect within seconds and cannot be undone. --- ## SDK & API Quickstart Source: https://aceteam.ai/docs/sdks Manage sandboxes, GPU compute, and ACET tokens from Python and TypeScript. AceTeam provides client SDKs for Python and TypeScript that wrap the [Hosted Instances](/docs/instances-overview) API -- managed containers with automatic safety enforcement. There is also a REST API for GPU compute and ACET token management. This guide covers both. **Hosted Instances vs. Citadel Sandboxes:** The SDKs manage hosted instances (containers with safety policies, cost tracking, and audit trails). For raw ephemeral containers on your own Citadel nodes, see the [Sandboxes](/docs/sandboxes) documentation instead. ## Sandbox SDKs The sandbox SDKs provide a high-level interface for managing hosted instances -- containers that run with automatic safety enforcement. ### Python SDK Install: ```bash pip install aceteam-sdk ``` Quick start: ```python import asyncio from aceteam_sdk import Sandbox async def main(): # Create and start a sandbox sandbox = await Sandbox.create( template="safeclaw", api_key="act_your_api_key", ) print(f"Sandbox {sandbox.id} is {sandbox.state}") # Fetch logs logs = await sandbox.logs(lines=50) for line in logs: print(line) # Stop and clean up await sandbox.destroy() asyncio.run(main()) ``` #### Authentication The SDK accepts an API key in two ways: - **Parameter**: `api_key="act_..."` on `Sandbox.create()` or any class method - **Environment variable**: `ACETEAM_API_KEY` The base URL defaults to `https://aceteam.ai`. Override with `base_url` or `ACETEAM_API_URL`. #### Creating Sandboxes ```python sandbox = await Sandbox.create( template="safeclaw", # Template ID name="my-agent", # Optional display name env_vars={"DEBUG": "1"}, # Optional env vars resource_limits={ # Optional resource limits "memory_mb": 2048, "cpu_millicores": 1000, }, ) ``` Create without auto-starting: ```python sandbox = await Sandbox.create(template="safeclaw", auto_start=False) await sandbox.start() ``` #### Context Manager The sandbox is destroyed automatically when the context exits: ```python async with await Sandbox.create(template="safeclaw") as sandbox: logs = await sandbox.logs() # sandbox is destroyed here ``` #### Lifecycle ```python await sandbox.start() # Start a created/stopped sandbox await sandbox.stop() # Stop the container await sandbox.restart() # Stop + start await sandbox.destroy() # Permanently delete ``` #### Inspect State ```python instance = await sandbox.refresh() # Re-fetch from API print(sandbox.state) # "created", "starting", "running", etc. print(sandbox.public_url) # Public URL if exposed ``` #### List and Retrieve ```python # List all sandboxes instances = await Sandbox.list(api_key="act_...") # Get a specific sandbox by ID sandbox = await Sandbox.get("uuid-here", api_key="act_...") # List available templates templates = await Sandbox.templates(api_key="act_...") ``` #### Error Handling ```python from aceteam_sdk import AceTeamAPIError try: sandbox = await Sandbox.get("nonexistent-id", api_key="act_...") except AceTeamAPIError as e: print(f"HTTP {e.status_code}: {e.detail}") ``` ### TypeScript SDK Install: ```bash npm install @aceteam/sandbox-sdk ``` Quick start: ```typescript import { Sandbox } from "@aceteam/sandbox-sdk"; const sandbox = await Sandbox.create({ template: "safeclaw" }); console.log(`Sandbox ${sandbox.id} is ${sandbox.state}`); const logs = await sandbox.logs(); console.log(logs); await sandbox.destroy(); ``` #### Authentication - **Parameter**: `apiKey: "act_..."` in the options object - **Environment variable**: `ACETEAM_API_KEY` Override the base URL with `baseUrl` or `ACETEAM_API_URL`. #### Creating Sandboxes ```typescript const sandbox = await Sandbox.create({ template: "safeclaw", name: "my-agent", envVars: { DEBUG: "1" }, resourceLimits: { memory_mb: 2048, cpu_millicores: 1000, }, }); ``` #### Lifecycle ```typescript await sandbox.start(); await sandbox.stop(); await sandbox.restart(); await sandbox.destroy(); ``` #### List and Retrieve ```typescript const instances = await Sandbox.list({ apiKey: "act_..." }); const sandbox = await Sandbox.get("uuid-here", { apiKey: "act_..." }); const templates = await Sandbox.templates({ apiKey: "act_..." }); ``` #### Error Handling ```typescript import { AceTeamAPIError } from "@aceteam/sandbox-sdk"; try { const sandbox = await Sandbox.get("nonexistent-id", { apiKey: "act_..." }); } catch (e) { if (e instanceof AceTeamAPIError) { console.error(`HTTP ${e.statusCode}: ${e.detail}`); } } ``` ### SDK API Coverage Both SDKs support the same operations: | Feature | Python | TypeScript | Status | | ---------------------- | --------------------------------- | --------------------------------- | ----------- | | Create | `Sandbox.create()` | `Sandbox.create()` | Available | | Start / Stop / Restart | `sandbox.start()` etc. | `sandbox.start()` etc. | Available | | Destroy | `sandbox.destroy()` | `sandbox.destroy()` | Available | | Get / List | `Sandbox.get()`, `Sandbox.list()` | `Sandbox.get()`, `Sandbox.list()` | Available | | Logs | `sandbox.logs()` | `sandbox.logs()` | Available | | Templates | `Sandbox.templates()` | `Sandbox.templates()` | Available | | Exec | `sandbox.exec()` | `sandbox.exec()` | Coming soon | | File operations | `sandbox.files.*` | `sandbox.files.*` | Coming soon | ## REST API: GPU Compute The GPU compute features are accessible through the REST API. All requests require a Bearer token. ### Provision a Model One API call to find a node and deploy a model: ```python import requests resp = requests.post( "https://aceteam.ai/api/fabric/provision", headers={"Authorization": "Bearer act_your_api_key"}, json={ "model": "meta-llama/Llama-3-8B-Instruct", "gpu_preference": "any", }, ) result = resp.json() print(f"Node: {result['node_id']}, Status: {result['status']}") ``` ```typescript const resp = await fetch("https://aceteam.ai/api/fabric/provision", { method: "POST", headers: { Authorization: "Bearer act_your_api_key", "Content-Type": "application/json", }, body: JSON.stringify({ model: "meta-llama/Llama-3-8B-Instruct", gpu_preference: "any", }), }); const result = await resp.json(); ``` ### GPU Metrics ```python resp = requests.get( f"https://aceteam.ai/api/fabric/nodes/{node_id}/metrics?period=24h", headers={"Authorization": "Bearer act_your_api_key"}, ) for point in resp.json()["data"][-3:]: print(f"GPU {point['gpu_util_pct']}%, " f"VRAM {point['vram_used_gb']}/{point['vram_total_gb']} GB") ``` ### Benchmarks ```python resp = requests.post( f"https://aceteam.ai/api/fabric/nodes/{node_id}/benchmark", headers={"Authorization": "Bearer act_your_api_key"}, json={"model_name": "llama3:8b", "num_iterations": 3}, ) agg = resp.json()["aggregate"]["tokens_per_sec"] print(f"Tokens/sec: median={agg['median']}, p95={agg['p95']}") ``` ### Node Reliability ```python resp = requests.get( f"https://aceteam.ai/api/fabric/nodes/{node_id}/reliability", headers={"Authorization": "Bearer act_your_api_key"}, ) r = resp.json() print(f"Reliability: {r['score']}/100 ({r['tier']})") ``` See the [GPU Compute reference](/docs/gpu-compute) for the full API surface including model caching, scale-to-zero, and power management. ## REST API: ACET Tokens ### Check Balance ```python resp = requests.get( "https://aceteam.ai/api/acet/balance", headers={"Authorization": "Bearer act_your_api_key"}, ) print(f"Balance: {resp.json()['balance']} ACET") ``` ### Purchase Tokens ```python resp = requests.post( "https://aceteam.ai/api/acet/checkout", headers={"Authorization": "Bearer act_your_api_key"}, json={"amount_acet": 10000}, # $10.00 ) print(f"Checkout: {resp.json()['checkout_url']}") ``` ### Self-Registration Create an organization and API key programmatically (no auth required): ```python resp = requests.post( "https://aceteam.ai/api/fabric/register", json={"email": "agent@example.com", "org_name": "My Org"}, ) creds = resp.json() # Use creds["api_key"] for subsequent requests ``` See the [ACET Tokens reference](/docs/acet-tokens) for details on pricing and the revenue split. ## Error Handling All REST endpoints return standard HTTP status codes with a `detail` field: | Status | Meaning | | ------- | ------------------------------------------------ | | 200/201 | Success | | 400 | Invalid request | | 401 | Missing or invalid API key | | 404 | Resource not found | | 429 | Rate limit exceeded (check `Retry-After` header) | | 502 | Node error | | 503 | Service unavailable | | 504 | Timeout | ## Next Steps - [Hosted Instances](/docs/instances-overview) -- managed containers with safety enforcement (what the SDKs wrap) - [GPU Compute](/docs/gpu-compute) -- model deployment, metrics, benchmarks, power management - [ACET Tokens](/docs/acet-tokens) -- billing, balance, self-registration - [Sandboxes](/docs/sandboxes) -- ephemeral containers on your Citadel nodes (separate from hosted instances) --- ## Endpoints Overview Source: https://aceteam.ai/docs/endpoints-overview Explore available API endpoint groups and request/response formats. # Endpoints Overview The AceTeam API provides programmatic access to agents, workflows, documents, and workbenches. This page describes the available endpoint groups, request and response formats, and conventions used across the API. ## Base URL All API requests should be made to: ``` https://aceteam.ai/api/v1 ``` All endpoints require HTTPS. Plain HTTP requests are rejected. ## Endpoint Groups ### Agents Manage and interact with your AI agents. | Method | Endpoint | Description | | ------ | ------------------ | ------------------------------------- | | `GET` | `/agents` | List all agents in your organization | | `GET` | `/agents/:id` | Get a specific agent by ID | | `POST` | `/agents/:id/chat` | Send a message and receive a response | ### Workflows Build and execute multi-step automation workflows. | Method | Endpoint | Description | | ------ | -------------------- | ---------------------------------- | | `GET` | `/workflows` | List all workflows | | `GET` | `/workflows/:id` | Get a specific workflow by ID | | `POST` | `/workflows/:id/run` | Execute a workflow with input data | ### Documents Upload and manage documents for agent knowledge bases. | Method | Endpoint | Description | | ------ | ---------------- | --------------------- | | `GET` | `/documents` | List all documents | | `GET` | `/documents/:id` | Get document metadata | | `POST` | `/documents` | Upload a new document | ### Workbenches Access agent testing and development environments. | Method | Endpoint | Description | | ------ | ------------------ | ------------------------ | | `GET` | `/workbenches` | List all workbenches | | `GET` | `/workbenches/:id` | Get a specific workbench | ### Gateway OpenAI-compatible reverse proxy with cost tracking, safety detection, and model routing. See the full [Gateway API documentation](/docs/gateway). | Method | Endpoint | Description | | ------ | ------------------------------ | ----------------------- | | `POST` | `/gateway/v1/chat/completions` | Proxied chat completion | ## Request Format All request bodies must be JSON. Include the `Content-Type` header: ```bash curl -X POST https://aceteam.ai/api/v1/agents/abc123/chat \ -H "Authorization: Bearer act_your_key_here" \ -H "Content-Type: application/json" \ -d '{"message": "Summarize the Q3 report."}' ``` ## Response Format All responses return JSON with a consistent structure. Successful responses include a `data` field: ```json { "data": { "id": "agt_abc123", "name": "Research Assistant", "status": "active" } } ``` List endpoints return an array under `data` along with pagination metadata: ```json { "data": [ ... ], "pagination": { "total": 42, "limit": 20, "offset": 0 } } ``` ## Error Responses Errors return an appropriate HTTP status code and a JSON body with an `error` field: ```json { "error": { "code": "not_found", "message": "Agent with ID 'agt_xyz' was not found." } } ``` | Status Code | Meaning | | ----------- | -------------------------------------------------------- | | `400` | Bad Request -- invalid parameters or malformed JSON | | `401` | Unauthorized -- missing or invalid API key | | `403` | Forbidden -- key lacks permission for this resource | | `404` | Not Found -- resource does not exist | | `429` | Too Many Requests -- rate limit exceeded | | `500` | Internal Server Error -- something went wrong on our end | ## Pagination List endpoints accept `limit` and `offset` query parameters: ``` GET /v1/agents?limit=20&offset=40 ``` - `limit` controls the maximum number of items returned (default: 20, maximum: 100). - `offset` controls how many items to skip from the beginning of the result set. ## MCP Integration All API endpoints are also available as MCP (Model Context Protocol) tools. This means you can connect AceTeam to Claude Code, Claude Desktop, and other MCP-compatible clients. Once configured, your agents, workflows, and documents are accessible directly from your development environment. See the [Technical FAQs](/docs/technical-faqs) for setup instructions. --- ## Gateway API Source: https://aceteam.ai/docs/gateway Route any OpenAI-compatible client through the AceTeam gateway for cost tracking, safety detection, model routing, and failover. # Gateway API The AceTeam gateway is an OpenAI-compatible reverse proxy. Point any client that speaks the OpenAI API at `https://aceteam.ai/api/gateway/v1` and authenticate with an `act_` API key. Every call gets cost tracking, safety detection, smart model routing, and automatic failover: no code changes beyond the base URL and key. If you're looking for the self-hosted version, see [AEP Safety Proxy](/docs/safety-proxy). The hosted gateway runs the same safety pipeline with additional features: multi-provider routing, org-level cost tracking, and credit-based billing. ## Quickstart ### 1. Create an API key Go to [API Keys](/api-keys) and click **Create API Key**. Copy the key: it starts with `act_` and is only shown once. ### 2. Make a request Any OpenAI-compatible client works. Set the base URL and API key: **curl** ```bash curl https://aceteam.ai/api/gateway/v1/chat/completions \ -H "Authorization: Bearer act_xxxx..." \ -H "Content-Type: application/json" \ -d '{ "model": "claude-haiku-4-5-20251001", "messages": [{"role": "user", "content": "Hello"}] }' ``` **Python** ```python from openai import OpenAI client = OpenAI( base_url="https://aceteam.ai/api/gateway/v1", api_key="act_xxxx...", ) response = client.chat.completions.create( model="claude-haiku-4-5-20251001", messages=[{"role": "user", "content": "Hello"}], ) print(response.choices[0].message.content) ``` **TypeScript / Node.js** ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://aceteam.ai/api/gateway/v1", apiKey: "act_xxxx...", }); const response = await client.chat.completions.create({ model: "claude-haiku-4-5-20251001", messages: [{ role: "user", content: "Hello" }], }); console.log(response.choices[0].message.content); ``` **Environment variables**: alternatively, set these and your client picks them up automatically: ```bash export OPENAI_BASE_URL=https://aceteam.ai/api/gateway/v1 export OPENAI_API_KEY=act_xxxx... ``` ### 3. View your dashboard See per-key cost tracking and safety signals at [aceteam.ai/gateway](/gateway). ## Models The gateway accepts standard model slugs (e.g., `claude-haiku-4-5-20251001`, `gpt-4o`, `gemini-2.0-flash`). The route resolver picks the cheapest available channel automatically based on your org's routing policy. You don't need to configure API keys for each provider. The gateway uses platform keys by default. Organizations can also bring their own keys (BYOK). Contact your admin to configure provider credentials. ## Routing policies Org admins can configure how the gateway picks a provider channel for each model. Set your routing policy at [Settings → Organization → Routes](/settings/organization/routes). | Policy | Behavior | | --------------------------- | -------------------------------------------------------------------------------------------- | | **Prefer Direct** (default) | Use the model's default channel if available, otherwise pick the cheapest alternative | | **Cheapest** | Always pick the channel with the lowest price per token | | **Pinned** | Use only your org's pinned channel; fail if unavailable | | **Balanced** | Cheapest within the default channel's priority tier, then remaining channels sorted by price | | **Failover Only** | Use the default channel only; never fall back to alternatives | Routing is per-model: you can pin one model to a specific provider while letting others route cheaply. ## Streaming Pass `"stream": true` in your request body. The gateway streams SSE chunks to your client in real time, identical to the OpenAI streaming format. Safety detectors run on the complete response after the stream finishes. If the post-stream safety check triggers a block, the gateway appends an `aep_safety_block` SSE event to the stream. ## Response headers Every response includes these headers: | Header | Description | Example | | ---------------------- | ---------------------------------------- | -------------------------- | | `X-AEP-Cost` | USD cost of the call | `0.00231` | | `X-AEP-Enforcement` | Safety decision | `pass`, `flag`, or `block` | | `X-AEP-Call-ID` | Unique call identifier (8-char hex) | `a1b2c3d4` | | `X-AEP-Classification` | Data sensitivity level | `public` | | `X-AEP-Flag-Reason` | Reason string (only on `flag` decisions) | `pii_detected` | | `X-AEP-Trace-ID` | Trace ID (only if sent in request) | `abc123` | ### Optional request headers You can send these headers to add governance context to a call: | Header | Purpose | Example | | ---------------------- | -------------------------------- | ------------------------- | | `X-AEP-Entity` | Identify who is making the call | `org:acme-corp` | | `X-AEP-Classification` | Data sensitivity level | `confidential` | | `X-AEP-Consent` | Data governance directives | `training=no,sharing=org` | | `X-AEP-Budget` | Max spend for this call (USD) | `0.50` | | `X-AEP-Trace-ID` | Link to a parent execution trace | `trace-abc123` | | `X-AEP-Sources` | Declare data sources in context | `doc:contract-123` | See [AEP Protocol Overview](/docs/aep-overview) for the full governance protocol. ## Error handling | Status | Code | When | | ------ | --------------------- | ------------------------------------------------------------------------------- | | `401` | `UNAUTHORIZED` | Missing or invalid `act_` API key | | `402` | `credits_exhausted` | Org has zero credits remaining | | `400` | `safety_block` | Request or response blocked by safety policy | | `429` | `budget_exceeded` | Call exceeds the `X-AEP-Budget` limit | | `503` | `gateway_unavailable` | Python backend is unreachable | | `503` | `fail_closed` | Safety check unavailable and policy requires it (includes `Retry-After` header) | Error responses follow the OpenAI error format: ```json { "error": { "message": "AEP safety: request blocked — pii_detected", "type": "aep_safety_block", "code": "safety_block" } } ``` ## Safety The gateway runs three built-in detectors on every call by default: - **PII Detection**: catches SSNs, emails, phone numbers, credit cards, and other PII. Severity: `high` → blocked. - **Agent Threat Detection**: detects dangerous system calls (port scanning, reverse shells, subprocess execution). Severity: `high` → blocked. - **Cost Anomaly**: flags calls costing 5x the session average. Severity: `medium` → flagged. Safety can be toggled per-session from the [gateway dashboard](/gateway). For more details, see [Safety Detectors](/docs/safety-detectors). ## Framework compatibility The gateway works with any client that calls the OpenAI-compatible API: | Framework | Configuration | | ------------------------ | -------------------------------------- | | **OpenAI Python/JS SDK** | Set `base_url` / `baseURL` | | **LangChain** | Set `openai_api_base` on the LLM | | **CrewAI** | Set `OPENAI_BASE_URL` env var | | **DSPy** | Set `api_base` on `dspy.LM` | | **Vercel AI SDK** | Use `createOpenAI({ baseURL: "..." })` | | **curl** | Use the full URL directly | ## Gateways vs. Sessions The [Gateways dashboard](/gateways) lets you create **sessions**: isolated gateway instances, each with its own API key, safety policy, and spending limit. Sessions are useful when you want: - Separate cost tracking per project or client - Different safety policies for different use cases - Spending caps on individual integrations The `act_` key from Settings → API Keys works globally for your org. Session-specific keys scope cost and policy to that session only. ## Next steps - [In-app quickstart](/gateway/quickstart): interactive 60-second setup - [AEP Safety Proxy](/docs/safety-proxy): self-hosted version for local development - [Safety Detectors](/docs/safety-detectors): detector architecture and custom detectors - [AEP Protocol Overview](/docs/aep-overview): the full accountability protocol --- ## AceTeam Whitepaper Source: https://aceteam.ai/docs/accountable-compute-whitepaper Sovereignty, cost transparency, and data governance in the agent economy. Why AI infrastructure needs an accountability layer, and how AceTeam built one. # The Accountability Layer for the Agent Economy ## Sovereignty, Cost Transparency, and Data Governance ## Executive Summary The world is deploying AI agents faster than it can account for them. Enterprises are building multi-step agent workflows that cross organizational boundaries, invoke external models, and process sensitive data, with no wire-level mechanism for tracking what it cost, what data was touched, or which organizations were involved. The infrastructure layer beneath these agents (Kubernetes, cloud APIs, GPU clusters) solves scheduling and scaling. It does not solve accountability. This gap is not a feature request. It is a structural absence in the AI infrastructure stack. Today's platforms can tell you that a GPU was utilized for 47 minutes. They cannot tell you that Agent A called Agent B, which called Agent C across three organizations, spent $47.30 across two model providers, cited 14 source documents, and kept PHI data within HIPAA-consented boundaries throughout. Accountable AI compute is the missing layer: open wire protocols that attach cost trees, citation chains, and data governance audit trails to every agent interaction, running on infrastructure the organization controls. This whitepaper examines why this layer is necessary, what it requires, and how organizations can adopt it incrementally without replacing their existing infrastructure. AceTeam has built this layer. The Agentic Execution Protocol (AEP) and Agent Compute Protocol (ACP) are open wire protocols: implemented in TypeScript and Python, deployed in production, and running real workloads for government, healthcare, and data center customers today. What follows is the framework behind them. ## A Concrete Example: Clinical Document Processing Before examining the framework, consider a real scenario that illustrates the accountability gap. A healthcare organization deploys an AI workflow for clinical document processing. A care coordinator uploads patient intake forms, referral letters, and lab results. The workflow executes: 1. **Document extraction**: structured data is pulled from PDFs and scanned forms using an OCR + language model pipeline 2. **Entity resolution**: patient identifiers, medications, diagnoses, and provider names are identified with confidence scores 3. **Summary generation**: a language model drafts a clinical summary, citing specific source documents for each finding 4. **PHI detection**: a separate model scans the output and flags protected health information for review before sharing 5. **Confidence calibration**: the system documents what it _doesn't_ know: "Referral mentions prior imaging but no results were included in the submitted documents" This workflow touches sensitive patient data, invokes multiple AI models, and produces a document used for clinical decision-making. Now ask the three accountability questions: **What did it cost?** The workflow invoked four models across two providers. The OCR ran on a local GPU; the summary generation used a cloud API. The total cost is split across infrastructure the healthcare organization owns and external services it pays for. Without cost attribution, the per-document cost is unknowable, and the organization cannot budget for scaling to 10,000 records per month. **Based on what?** The clinical summary states that the patient has a history of hypertension based on intake documentation. Which source document contained that finding? Which model extracted it? What was the confidence score? If a lab result is later corrected, which other conclusions in the summary depended on it? Without citation chains, the summary's clinical integrity is unverifiable. **Who saw the data?** Patient records contain PHI protected under HIPAA. Which organizations processed this data? Did the cloud provider have a BAA authorizing PHI handling? Was any data transmitted outside the required jurisdiction? Without governance audit trails, the healthcare organization cannot demonstrate chain of custody to a regulator. This is not a hypothetical. These are the questions that procurement officers, compliance teams, and regulators ask before approving AI deployment in healthcare. And this pattern (multi-model workflows on sensitive data, requiring cost attribution, source tracing, and chain of custody) repeats in every regulated industry: financial compliance, legal discovery, government casework, insurance claims processing. The specifics change; the accountability requirements do not. Accountable AI compute answers all three questions: structurally, automatically, for every workflow execution. ## The Accountability Imperative ### Why Now Three forces are converging to make AI accountability a boardroom priority. **AI agents are becoming autonomous.** The shift from copilots (human-in-the-loop) to agents (human-on-the-loop) means AI systems are making decisions, calling external services, and processing data without human review of each step. When a copilot hallucinates, a human catches it. When an agent hallucinates inside a 12-step workflow at 2 AM, nobody does. **Multi-organization workflows are the norm, not the exception.** A consulting firm's strategy agent calls a research firm's analysis agent, which calls a data processing service running on a client's GPU cluster. Three organizations, three runtimes, one workflow. Today's infrastructure has no mechanism to track costs, data movement, or decision provenance across these boundaries. **Regulation is catching up.** The EU AI Act requires transparency and traceability for high-risk AI systems. Canada's PIPEDA and proposed Bill C-27 impose strict data residency requirements. The US CLOUD Act creates jurisdictional conflicts for data stored on foreign-owned cloud infrastructure. Organizations deploying AI agents without accountability infrastructure are accumulating compliance debt with every workflow execution. With enterprise AI agent spending projected at $71B by 2028, the scale of unaccountable compute is growing faster than the tools to govern it. ### The Four Dimensions of AI Accountability Accountability in AI compute spans four interdependent dimensions. Addressing one without the others creates blind spots. **Cost accountability** answers: _What did this workflow cost, broken down by step, model, and organization?_ Current AI infrastructure provides aggregate billing: monthly invoices from OpenAI, AWS bills by service. But when a single workflow invokes multiple models across multiple providers, the per-workflow cost is unknowable. Budget enforcement is impossible. Chargebacks to internal departments or external clients require manual reconciliation. Cost accountability means every workflow step produces a cost record. Cost records compose into cost trees: hierarchical structures where parent tasks aggregate the costs of their children. An organization can see that a document processing workflow cost $12.40: $3.20 for extraction (GPT-4o), $1.80 for classification (Claude), $7.40 for 15 minutes of GPU compute on their own hardware. **Source attribution** answers: _What data informed this AI output, through which models, with what confidence?_ AI outputs without provenance are opinions, not analysis. When a research agent concludes that a drug pipeline is promising, the citation chain should trace that conclusion through the intermediate analysis steps back to the original clinical data, patent filings, and market reports. If any source is retracted or updated, every downstream conclusion that depended on it can be identified. Source attribution is not a reporting feature. It is a structural property of the execution protocol: citations propagate automatically as results flow up through the workflow. **Data governance** answers: _Which organizations processed which data, with what consent, under which jurisdiction?_ In multi-organization workflows, data crosses organizational boundaries at every agent-to-agent handoff. A healthcare workflow might send patient data to an extraction service, which sends anonymized results to an analysis service, which returns conclusions to the originating hospital. Each boundary crossing requires consent verification, and the complete data flow must be auditable after the fact. Data governance in accountable compute means governance policies travel with the data. Every agent in the workflow can verify: does this organization have consent to process this data classification? If not, the workflow halts before the violation occurs, not after. **Operational sovereignty** answers: _Where did this workflow execute, on whose hardware, under whose control?_ Sovereignty is not just about data residency. It is about operational authority: the ability to run AI workloads on infrastructure you control, inspect the execution environment, and verify that no external party had access. For government agencies processing classified information, healthcare organizations handling PHI, or financial institutions managing trading strategies, this is non-negotiable. Operational sovereignty requires that the compute fabric, the actual hardware running the workloads, can be deployed on premises, in a private data center, or on dedicated hardware within a cloud region. The orchestration layer must work identically regardless of where the hardware sits. ## The Infrastructure Gap ### What Existing Tools Solve The current AI infrastructure stack is mature and capable, for scheduling and scaling. It is not designed for accountability. **Kubernetes** schedules containers on nodes. It manages pod lifecycle, health checks, and resource allocation. It does not know that the container running on Node 3 is processing data that requires HIPAA consent, or that the workflow it belongs to has a $50 budget cap. **Cloud AI platforms** (AWS Bedrock, Azure AI, GCP Vertex) provide managed model access with per-request billing. They do not provide cost attribution across multi-step workflows, citation chains between model invocations, or governance enforcement at organizational boundaries. **GPU compute platforms** (CoreWeave, Modal, Lambda) provide raw or serverless GPU access. They bill per GPU-hour. They do not provide agent-initiated resource negotiation, budget enforcement before allocation, or usage reports that feed into workflow-level cost trees. **Workflow engines** (n8n, Airflow, Temporal) orchestrate multi-step processes. They track step completion and handle retries. They do not produce cost records, citation chains, or governance audit trails as structural outputs of execution. **Observability platforms** (Datadog, Grafana, LangFuse) monitor infrastructure and application metrics. They provide dashboards and alerts. They do not produce compliance-grade evidence that a specific workflow executed within budget, cited its sources, and respected data governance boundaries. Each of these tools excels at its purpose. The gap is not in any single tool. It is in the absence of a protocol layer that connects them into an accountable whole. ### What Accountable Compute Requires Accountable AI compute is not a product category. It is an infrastructure layer: a set of wire protocols that any platform, any agent framework, and any compute provider can implement. It requires two complementary protocols. **An execution protocol** that attaches accountability metadata to every agent interaction. Context (budget, governance rules, tracing identifiers) flows downstream when an agent calls a sub-agent. Results (cost records, citations, audit trails) flow upstream when the sub-agent returns. The protocol is structural: accountability happens because the wire format requires it, not because a developer remembered to log it. **A compute protocol** that enables agents to negotiate resources dynamically. Instead of pre-provisioned infrastructure, agents request workers at runtime: "I need 5 GPU workers for 10 minutes, budget $15." A resource broker verifies capacity, enforces budget constraints, returns a queue endpoint, and produces a usage report on release. The compute protocol connects to the execution protocol: compute costs feed into the workflow's cost tree automatically. Together, these protocols create a trust layer that enables AI agents to transact across organizational boundaries. An agent can call an external service knowing that costs will be tracked, data governance will be enforced, and an audit trail will exist, without either party implementing custom integration logic. ### Why Not Just Log Everything? The most common response to the accountability gap is: "We already have observability tools. We can log everything and reconstruct accountability after the fact." This approach fails for three reasons. **Logs are internal; protocols are interoperable.** An organization's Datadog logs capture what happened inside their infrastructure. They do not capture what happened inside a partner organization's infrastructure when a cross-org workflow executed. Accountability across organizational boundaries requires a shared wire format, not shared log access. **Logs are retrospective; governance must be preventive.** Logs tell you that PHI data was sent to an unauthorized processor, after the violation occurred. Protocol-level governance halts the workflow before the data crosses the boundary. The difference is between a compliance incident and a compliance near-miss. Regulators care about the distinction. **Logs are unstructured; cost trees are compositional.** Reconstructing a per-workflow cost breakdown from Kubernetes pod metrics, cloud API billing, and application logs requires custom ETL for every combination of infrastructure. A wire protocol that produces cost records at every step composes them automatically. The cost tree is a structural output, not an analytical exercise. Observability tools are essential for operational monitoring. They are not a substitute for protocol-level accountability. ## A Framework for Accountable AI Compute ### Design Principles Accountable AI compute should follow five principles that enable adoption without requiring organizations to replace their existing infrastructure. **Protocol, not product.** Accountability should be defined as an open wire protocol, not locked inside a proprietary platform. Any agent framework (LangChain, CrewAI, AutoGen) should be able to produce and consume accountability metadata using standard HTTP/JSON or gRPC. The protocol specifications and core implementations should be open source, Apache 2.0 or equivalent, so that adoption is not gated by a single vendor. Vendor lock-in defeats the purpose of sovereignty. **Structural, not optional.** Accountability must be a property of the execution format, not a feature flag. Every workflow execution should produce cost records and audit trails by default. If accountability is optional, it will be omitted under deadline pressure, exactly when it matters most. **Incremental adoption.** Organizations should be able to adopt accountability at their own pace. A minimal implementation tracks costs. A fuller implementation adds citation chains. A complete implementation adds data governance with jurisdictional enforcement. Each level provides standalone value. **Context down, results up.** In multi-agent workflows, execution context (budget, governance rules, tracing) propagates downstream from caller to callee. Results (cost records, citations, audit trails) propagate upstream from callee to caller. This mirrors the natural structure of delegation and is the principle that makes accountability composable across organizational boundaries. **Infrastructure agnostic.** The protocols must work on cloud, on premises, on bare metal, and in hybrid environments. An organization running Kubernetes on AWS and another running K3s in their own data center should produce identical accountability metadata. ### Conformance Levels Not every organization needs full accountability from day one. A tiered conformance model allows incremental adoption. | Level | Name | What it provides | Who needs it | | ----- | --------------- | ----------------------------------------------------- | ----------------------------------------------------------- | | 1 | **Minimal** | Unique execution IDs and timing | Any organization wanting basic traceability | | 2 | **Traceable** | Cost records with per-step attribution | Organizations needing budget visibility and chargebacks | | 3 | **Accountable** | Citation chains linking outputs to source data | Regulated industries, research organizations, legal | | 4 | **Governed** | Data governance enforcement with consent verification | Healthcare (HIPAA), government (PIPEDA), financial services | Each level includes all capabilities of the levels below it. An organization at Level 2 automatically gets Level 1 traceability. Moving from Level 2 to Level 3 requires adding citation metadata to workflow steps. It does not require re-architecting the system. ## Sovereign Compute: Accountability Meets Infrastructure ### The Sovereignty Paradox Many organizations adopting "sovereign AI" face a structural irony: the platform promising sovereignty runs on infrastructure they don't control. A SaaS platform hosted on AWS that offers "sovereign AI deployment" is sovereign in the data plane (workloads run on customer hardware) but not in the control plane (orchestration runs on Amazon's infrastructure). This is not inherently wrong. Tailscale's control plane is cloud-hosted while its tunnels are peer-to-peer, and this is an accepted architecture. But for the most sensitive use cases (classified government data, healthcare PHI, financial trading strategies), the control plane itself must be self-hostable. True operational sovereignty requires three properties: 1. **Data plane sovereignty**: workloads execute on hardware the organization owns or controls 2. **Control plane optionality**: the orchestration layer can be cloud-hosted for convenience or self-hosted for maximum control 3. **Network sovereignty**: communication between nodes uses encrypted tunnels that don't traverse third-party infrastructure ### Dynamic Compute Allocation Traditional cloud computing uses static provisioning: deploy N workers, scale based on metrics, pay for idle capacity. AI agent workloads are bursty and unpredictable: a workflow might need 20 GPU workers for 3 minutes, then nothing for an hour. Agent-initiated compute allocation inverts this model. The agent tells the infrastructure what it needs: > "I need 5 GPU workers with B300 capability for 10 minutes. My budget is $15. The data is classified as PIPEDA-regulated and must stay in Canadian jurisdiction." The resource broker: - Verifies capacity exists on nodes matching the capability and jurisdiction requirements - Checks budget against current pricing - Reserves workers atomically (no over-commitment) - Returns a queue endpoint for job submission - Produces a usage report on release that feeds into the workflow's cost tree This model has three advantages over static provisioning: **Cost alignment.** Organizations pay for compute they actually use, measured in worker-minutes, not idle capacity. The usage report connects directly to the cost tree, enabling accurate chargebacks. **Governance enforcement.** The allocation request includes data classification and jurisdiction requirements. The broker only allocates workers on nodes that satisfy these constraints. Governance is enforced at the infrastructure level, not the application level. **Multi-tenant efficiency.** Multiple organizations can share the same physical infrastructure while maintaining strict isolation. Each organization's agents see only the capacity allocated to them. The broker manages the shared pool. ## Industry Implications ### Government and Public Sector Government agencies face the most stringent sovereignty requirements. Citizen data must remain within national borders. AI decision-making must be auditable. Vendors must demonstrate compliance, not just claim it. Accountable AI compute provides automatic evidence generation: every workflow execution produces a verifiable record of what data was processed, on which hardware, in which jurisdiction, at what cost. This transforms compliance from a manual audit exercise into a structural property of the infrastructure. For agencies processing classified information or sensitive government data, the ability to run the complete stack (orchestration, compute, and storage) on government-controlled hardware eliminates the jurisdictional risks inherent in cloud-hosted platforms. ### Healthcare Healthcare organizations process some of the most sensitive data in existence. Patient records, genomic data, and clinical trial results require HIPAA compliance in the US and equivalent protections globally. When AI workflows process this data across organizational boundaries, a hospital's diagnostic agent calling a research institution's analysis service, every boundary crossing must verify consent. Data governance in accountable compute enforces consent at the protocol level. Governance policies travel with the data, and every agent in the workflow verifies classification and consent before processing. If consent is missing, the workflow halts before the violation occurs, creating a provable record that protected data was never exposed to unauthorized processors. ### Financial Services Financial institutions use AI for fraud detection, risk modeling, trading strategies, and compliance monitoring. These workloads require both performance (low-latency inference) and accountability (auditability of AI-driven decisions). When a trading algorithm recommends a position based on AI analysis, regulators require the ability to trace that recommendation back through the models, data sources, and decision logic that produced it. Citation chains in accountable compute provide this traceability automatically. Every AI output carries its provenance: which models processed which data to produce which intermediate results. This is not a log to be assembled after the fact; it is a structural output of execution. ### Data Center Operators Data center operators with GPU infrastructure face a commoditization problem. Raw GPU rental is a race to the bottom: a B300 GPU rents for approximately $4.75/hour, and that price only decreases as supply grows. A rack of 24 GPUs generates roughly $998K/year in raw rental revenue. Margins are thin and defensibility is zero; customers switch to whoever is cheapest. Wrapping compute in accountable, sovereign, multi-tenant infrastructure transforms the economics entirely. Operators have two pricing models available: **Managed compute markup.** The same GPU at $4.75/hr raw becomes $10-15/hr as managed sovereign compute: with per-tenant isolation, compliance evidence generation, SLA guarantees, and governance enforcement included. The same 24-GPU rack generates $2.1-3.1M/year. **Per-workflow pricing.** The operator decouples pricing from GPU-hours entirely: $0.50 per document processed, $2 per report generated, $5 per compliance audit produced. The customer pays for outcomes, not infrastructure. The operator captures the efficiency gains from batching, caching, and multi-tenant scheduling. Both models can coexist. The platform takes a 10-15% brokerage fee on compute transactions, aligning incentives: the platform earns more when operators serve more customers, and operators earn more when the platform sends them workloads. The pitch to a data center operator is no longer "rent your GPUs." It is: "Turn your GPUs into a managed AI platform with per-workflow billing, compliance evidence, and multi-tenant isolation, without building the software yourself." That is a fundamentally different business from commodity GPU rental, and it commands fundamentally different margins. ## Implementation Considerations ### Starting Points Organizations considering accountable AI compute should begin with the dimension most relevant to their immediate needs. **If cost visibility is the priority:** Start with cost records. Instrument existing workflows to produce per-step cost attribution. This provides immediate value (accurate chargebacks, budget forecasting) and establishes the foundation for fuller accountability. **If compliance is the priority:** Start with data governance. Map data classifications to workflow steps. Define consent requirements at organizational boundaries. This addresses the most urgent regulatory exposure. **If sovereignty is the priority:** Start with infrastructure. Deploy compute on controlled hardware. Establish encrypted mesh networking. Verify data residency. This resolves the foundational question of where workloads execute before adding accountability metadata. ### Integration With Existing Infrastructure Accountable compute is an additive layer, not a replacement. Organizations should expect to: - **Keep their existing orchestration** (Kubernetes, cloud platforms): accountable compute adds protocols on top - **Keep their existing AI frameworks** (LangChain, custom agents): the protocols integrate via standard HTTP/JSON - **Keep their existing monitoring** (Prometheus, Datadog): accountability metadata complements operational metrics - **Add protocol-level instrumentation**: the execution and compute protocols produce structured metadata that feeds into existing dashboards and audit systems ### Maturity Model | Maturity | Characteristics | Typical timeline | | ------------- | --------------------------------------------------------------------------------------------- | -------------------------------------------------- | | **Reactive** | No cost tracking. No audit trail. Compliance is manual. | Starting point for most organizations | | **Visible** | Per-workflow cost attribution. Basic execution logging. | 4-8 weeks of instrumentation | | **Traceable** | Citation chains. Source attribution for AI outputs. | 8-16 weeks, requires workflow modification | | **Governed** | Data classification enforcement. Consent verification at boundaries. | 16-24 weeks, requires governance policy definition | | **Sovereign** | Full stack on controlled hardware. Self-hosted control plane. Compliance evidence generation. | 6-12 months, requires infrastructure investment | Organizations do not need to reach "Sovereign" maturity to derive value. Each level provides standalone benefits. The journey is incremental. ## Key Takeaways **The accountability gap is structural, not incidental.** Today's AI infrastructure was designed for scheduling and scaling, not for tracking costs, provenance, and data governance across organizational boundaries. Patching this with logging and dashboards treats the symptom, not the cause. **Accountability must be protocol-level.** Just as HTTPS made encryption structural (not optional) in web communication, accountable compute must make cost tracking, citation chains, and governance enforcement structural in AI execution. When accountability is a wire protocol, it happens automatically, regardless of which framework, which model, or which hardware is used. **Sovereignty is the foundation, not the ceiling.** Organizations that control their compute infrastructure have the ability to enforce accountability. Organizations that depend on third-party infrastructure must trust that the provider enforces it on their behalf. For regulated industries, this distinction is the difference between demonstrable compliance and assumed compliance. **The agent economy needs infrastructure, not just agents.** The bottleneck in AI adoption is shifting from creation (building agents) to operation (running them accountably at scale). The organizations that build accountability into their AI infrastructure now will have a structural advantage as multi-agent workflows become the norm, not because they can build agents faster, but because they can prove what those agents did. **Start now, start incrementally.** Full accountable compute is a journey, not a switch. Begin with cost visibility. Add citation chains. Layer in governance. Each step provides immediate value and reduces the compliance debt that accumulates with every unaccountable workflow execution. ## Next Steps **Read the full protocol specification.** The [Agentic Execution Protocol (AEP) whitepaper](/docs/aep-whitepaper) defines the wire format for cost trees, citation chains, and data governance in multi-agent workflows. The [ACP overview](/docs/acp-overview) covers dynamic compute allocation. Both protocols are open source (Apache 2.0). **Start a conversation.** If your organization is evaluating sovereign AI infrastructure, accountable compute for regulated workloads, or data center platform enablement, we'd welcome the opportunity to discuss your requirements. Reach us at [jason@aceteam.ai](mailto:jason@aceteam.ai). **See it in action.** AceTeam's platform implements both protocols today: visual workflow builder, sovereign compute fabric, and accountability metadata on every execution. [Request a demo](https://aceteam.ai) to see accountable AI compute running on real infrastructure. **For investors.** Read about [AceTeam's vision, traction, and team](/docs/deck). --- _AceTeam.ai builds accountable AI compute infrastructure. The Agentic Execution Protocol (AEP) and Agent Compute Protocol (ACP) are open wire protocols for cost tracking, source attribution, data governance, and dynamic compute allocation in multi-agent workflows._ _Copyright 2026 AceTeam.ai. All rights reserved._ --- ## Strategic Overview Source: https://aceteam.ai/docs/strategic-overview AceTeam: The Operating System for Sovereign AI # AceTeam: A Strategic Overview ## The Operating System for Sovereign AI ## The Sovereignty Imperative ### Nations and critical industries face a crisis of digital sovereignty. - **The Dilemma:** Every modern organization must use AI to remain competitive. However, regulations and national security concerns often forbid them from sending sensitive citizen, customer, or state data to foreign-owned cloud platforms. - **The Unacceptable Trade-Off:** This creates an impossible choice: sacrifice security to innovate, or sacrifice innovation to remain secure. - **The Result:** A multi-billion dollar market is paralyzed. Progress is stalled, not by a lack of technology, but by a lack of a trusted, sovereign platform to run it on. ## Our Solution: The Sovereign OS ### AceTeam provides the control plane for an organization's private AI infrastructure. We are not another AI model. We are the foundational software layer that allows organizations to run _any_ model on their own private hardware with the simplicity and power of a managed cloud. - **BUILD:** Use low-code visual tools to create sophisticated AI agents and workflows, securely connecting to internal data sources with full auditability. - **DEPLOY:** Install our lightweight agent (`Citadel`) on any server in minutes, transforming disconnected hardware into a unified, manageable AI fabric. - **GOVERN:** Manage all AI operations from a single dashboard. Ensure compliance, monitor resources, and guarantee that sensitive data never leaves the network perimeter. **We empower organizations to turn their existing hardware into a private, managed AI cloud.** ## BUILD: Low-Code AI Development ### Agent Management & Templates Pre-built agent templates for government, healthcare, and enterprise (from 911 dispatch to customer support) let teams deploy production-ready AI in minutes, not months. ![Agent Management: pre-built templates for government, healthcare, and enterprise](/images/pitch/rag-pipeline.png) ### AI Workbench An interactive development environment for testing, tuning, and iterating on AI agents in real time. Connect tools, select models, and validate behavior before deployment. ![AI Workbench: interactive agent development with tools and model selection](/images/pitch/workbench.png) ### Agent Builder Full control over agent configuration: system prompts, tool bindings, knowledge base connections, and dynamic variables, all from a single interface. ![Agent Builder: system prompt, tools, data sources, and variables configuration](/images/pitch/agent-builder.png) ### Visual Workflow Editor A drag-and-drop DAG editor for building multi-step AI workflows. Chain LLMs, data transforms, human review gates, and API calls into auditable pipelines. ![Visual Workflow Editor: DAG-based workflow for meeting summaries](/images/pitch/workflow-editor.png) ### Workflow Templates A library of ready-made workflow templates for common use cases (document processing, report generation, data extraction), accelerating time-to-value. ![Workflow Templates: library of pre-built workflow templates](/images/pitch/templates.png) ### Data Connectors & MCP Extensible integrations via the Model Context Protocol (MCP). Connect agents to internal databases, APIs, file systems, and third-party services without writing code. ![Data Connectors: MCP server integrations for databases, APIs, and services](/images/pitch/data-connectors.png) ### Voice Agents & Embeddable Chat Deploy AI agents as embeddable chat widgets or voice interfaces. Configure appearance, behavior, and branding, then embed with a single line of code. ![Voice Agents: embeddable chatbot configuration with live preview](/images/pitch/voice-agents.png) ## DEPLOY: Sovereign Compute Fabric ### Your hardware. Your network. Our management layer. Install the `Citadel` agent on any server in minutes. It joins the encrypted mesh network and becomes part of the organization's private AI fabric, fully managed from the AceTeam dashboard. No data leaves the perimeter. No cloud dependency. Full operational control. ![Sovereign Compute Fabric: network nodes, IPs, and status dashboard](/images/pitch/sovereign-compute.png) ## GOVERN: Compliance & Control ### Team & Permissions Role-based access control with granular permissions. Assign roles, set usage quotas, and audit every action across the organization. ![Team & Permissions: members, roles, quotas, and access control](/images/pitch/collaboration.png) ### Multi-Tenancy Full organizational isolation with seamless switching. Each tenant gets its own agents, workflows, data, and audit trail, managed from a single account. ![Multi-Tenancy: organization switcher with isolated environments](/images/pitch/multi-tenant.png) ## Validation: Government of Alberta ### Our platform was validated at the highest level of government. In a competitive Request for Proposal (RFP) with the Government of Alberta, AceTeam was selected as the technical winner over global incumbents and major consulting firms. - **The Mandate:** Provide the foundational AI platform for an entire government ministry. - **The Validation:** After a week-long, intensive technical evaluation, our platform was proven to be more capable, more secure, and more aligned with the government's sovereign requirements than any other solution on the market. - **The Takeaway:** This process provided an undeniable, third-party validation of our technology and our unique ability to meet the stringent demands of public sector clients. ## Solutions by Industry ### One platform, tailored for each vertical. AceTeam serves five core markets, each with dedicated templates, compliance controls, and deployment patterns: - **[Government](/solutions/government)**: Pre-built templates for zoning chatbots, document Q&A, citizen services, and data extraction. HIPAA, PIPEDA, GDPR, and SOC 2 compliance built in. Air-gapped deployment via Sovereign Compute Fabric. - **[Consultants](/solutions/consultants)**: Multi-client AI agents, automated deliverables, proposal generation, and knowledge management. Serve 15+ clients from one platform with full tenant isolation. - **[Enterprise](/solutions/enterprise)**: Sovereign deployment with SSO/OIDC, audit logs, compliance evidence, and dedicated support. AI operator training curriculum included. - **[Data Centers](/solutions/data-centers)**: Transform GPU infrastructure into a managed AI services business. Multi-tenant isolation, usage-based billing, tag-based GPU orchestration, and white-label options. - **[Education](/solutions/education)**: Hands-on AI training with gamified exercises and progressive curriculum. Currently running at San Jose State University. Student project templates and instructor dashboards. ## The Sovereign Cloud Ecosystem ### Technology alone is not enough. We are building the ecosystem to accelerate adoption. AceTeam is designed to be the software layer in a complete sovereign AI stack. We partner with national-scale infrastructure providers (data center operators, hardware vendors, and managed service providers) to offer turnkey sovereign cloud solutions. Our platform sits at the center of this ecosystem: partners provide the physical infrastructure and local expertise, AceTeam provides the operating system that ties it all together. The result is a clear, de-risked path for government and enterprise to modernize their AI infrastructure without compromising on sovereignty. ## Engage With Us ### The foundational platform for sovereign AI is here. AceTeam is partnering with public sector organizations, enterprises in regulated industries, and infrastructure providers who share our mission to build secure, sovereign AI infrastructure. **[Contact Us](/company/contact)** --- ## General FAQs Source: https://aceteam.ai/docs/general-faqs Common questions about pricing, models, privacy, and team management. # General FAQs Answers to the most common questions about the AceTeam platform, including available models, pricing, data privacy, and team management. ## What AI models are available? AceTeam integrates with leading AI providers so you can choose the best model for each task. Currently supported models include: - **OpenAI** -- GPT-4o, GPT-4o mini, and other GPT variants - **Anthropic** -- Claude Opus, Sonnet, and Haiku - **Google** -- Gemini Pro and Gemini Flash - **Deepseek** -- Deepseek reasoning and chat models New models are added regularly as providers release them. You can select which model an agent uses when configuring it in the Agent Builder. ## How does pricing work? AceTeam uses a credit-based pricing system. All resources (workflows, agents, API keys, webhooks, cron jobs, and files) are unlimited to create on every plan. Usage (running workflows, agent messages, voice calls, etc.) is deducted from your included monthly credits. Each plan includes a monthly credit grant: Free ($1/mo), Pro ($20/mo), Max ($100/mo). If you need more credits than your plan provides, you can purchase pay-as-you-go top-ups at any time without changing your plan. Credit costs vary by model and operation. For example, a GPT-4o conversation uses more credits than a GPT-4o mini conversation. Detailed pricing for each model and operation is displayed on the [Plans page](/plans). ## Is my data private? Yes. Your data is scoped to your organization and is never shared with other users or organizations. AceTeam does not use your data to train AI models. For organizations with stricter compliance requirements, Sovereign Compute allows you to run AI workloads on your own infrastructure. This means your data never leaves your network -- all processing happens on hardware you control. ## Can I use my own API keys? AceTeam provides built-in access to all supported AI models through the credit system, so you do not need your own API keys to get started. If you prefer to use your own model provider keys, Sovereign Compute lets you run models on your own hardware with your own credentials. This gives you full control over model access, costs, and data residency. ## How do I invite team members? 1. Navigate to **Settings > Members**. 2. Click **Invite**. 3. Enter the email address of the person you want to invite. 4. Select their role (Admin or Member). 5. Click **Send Invitation**. The invited person will receive an email with a link to join your organization. Once they accept, they will have access to the agents, workflows, and documents shared within the organization. ## What is the difference between plans? All plans allow unlimited creation of workflows, agents, API keys, webhooks, cron jobs, and files. Plans differ in the amount of included monthly credits, team size limits, feature access (BYOK, audit log retention, deployment options), and support level. The Free plan includes $1/mo in credits for individual experimentation, while paid plans provide larger credit grants, lower usage markup, and features like priority support and self-hosting. Visit the [Plans page](/plans) for a detailed comparison of all plans and their included features. ## Can I export my data? Yes. Agents and workflows can be exported from the platform in JSON format. To export an agent or workflow: 1. Open the agent or workflow you want to export. 2. Click the options menu (three dots). 3. Select **Export**. The exported file contains the full configuration, including system prompts, tool definitions, and workflow node structure. You can use exported files to back up your work, share configurations with other organizations, or import them into a different AceTeam workspace. Document files that you uploaded to knowledge bases can be re-downloaded from the Documents section at any time. --- ## Technical FAQs Source: https://aceteam.ai/docs/technical-faqs Answers about rate limits, file sizes, WebSockets, and self-hosting. # Technical FAQs Answers to common technical questions about rate limits, file handling, real-time communication, self-hosting, and integrations. ## What are the rate limits? Rate limits vary by plan and are enforced per API key. The current limits are included in the response headers of every API call: - `X-RateLimit-Limit` -- maximum requests per window - `X-RateLimit-Remaining` -- requests remaining in the current window - `X-RateLimit-Reset` -- Unix timestamp when the window resets If you exceed the limit, the API returns a `429 Too Many Requests` response. Implement exponential backoff in your client to handle this gracefully. If your workload consistently exceeds your plan's limits, consider upgrading or contacting support for a custom allocation. ## What file types can I upload? AceTeam supports a range of common file formats for document ingestion into agent knowledge bases: - **Text**: PDF, TXT, Markdown (.md) - **Data**: CSV, JSON, XLSX - **Images**: PNG, JPG, JPEG, WebP - **Code**: Most plain-text source files Uploaded documents are processed and indexed for retrieval-augmented generation (RAG), allowing your agents to reference their contents during conversations. ## What are the file size limits? File size limits depend on your plan: | Plan | Maximum File Size | | ---------- | ----------------- | | Free | 10 MB per file | | Pro | 50 MB per file | | Enterprise | Custom | If you need to process larger files, consider splitting them into smaller sections before uploading, or contact support to discuss Enterprise options. ## Does AceTeam support WebSockets? Yes. AceTeam uses WebSockets for real-time chat interactions with agents, providing a responsive streaming experience as the agent generates its response. WebRTC is used for real-time voice conversations, enabling low-latency two-way audio communication with voice-enabled agents. Both protocols are handled automatically by the AceTeam web interface. If you are building a custom integration, refer to the API documentation for WebSocket connection details. ## Can I self-host AceTeam? Sovereign Compute gives you control over where your AI workloads run by allowing you to connect your own hardware (GPU servers, databases, and other infrastructure) to the AceTeam platform. Your data stays on your network while you still benefit from the AceTeam interface and orchestration layer. AceTeam supports fully self-hosted deployment: the entire platform, including the control plane, can run on your own infrastructure with no external dependencies. ## How does the workflow engine handle errors? The workflow engine provides several mechanisms for handling errors during execution: - **Node-level retry**: Individual nodes can be configured with retry policies, including the number of attempts and delay between retries. - **Error propagation**: When a node fails and exhausts its retries, the error propagates to dependent nodes. You can add conditional branches to handle failures gracefully. - **Execution logs**: Every workflow run produces detailed execution logs that record the input, output, and status of each node. These logs are accessible from the workflow detail page and are useful for debugging failed runs. If a workflow fails, you can review the logs to identify the failing node, fix the issue, and re-run the workflow. ## What browsers are supported? AceTeam supports the latest two major versions of the following browsers: - Google Chrome - Mozilla Firefox - Apple Safari - Microsoft Edge For the best experience, keep your browser up to date. WebRTC-based voice features require a browser that supports the WebRTC standard, which all listed browsers do. ## How do I connect MCP tools to Claude Code? AceTeam exposes its agents, workflows, documents, and tools over the Model Context Protocol (MCP), so you can use them directly from Claude Code and other MCP-compatible clients. The short version: 1. Add the server, then start a new `claude` session: ```bash claude mcp add --transport http --scope user aceteam https://aceteam.ai/mcp ``` 2. Run `/mcp` inside Claude Code, pick **aceteam**, and complete sign-in in your browser. No API key is needed; authentication is OAuth. The older path `https://aceteam.ai/api/mcp/aceteam/mcp` still works as a legacy alias, so an existing server entry configured with it keeps working. For the full walkthrough, including verification and troubleshooting when tools do not show up, see [Connect AceTeam to Claude Code](/docs/connect-claude-code). --- ## Platform Overview Source: https://aceteam.ai/docs/platform-overview What AceTeam is, its key features, and who it's built for. AceTeam is an AI-driven platform for building, managing, and deploying AI agents and workflow automation. Whether you are a consultant automating client deliverables, a team lead streamlining internal processes, or a developer integrating AI into products, AceTeam gives you the tools to get work done with AI, without writing code. ## Key Features ### AI Agent Workbench The Workbench is where you create and test AI agents. Pick a model: OpenAI (GPT-5, GPT-4.1, o3, o4-mini), Anthropic Claude, Google Gemini, Deepseek, and more. Write a system prompt describing your agent's role, tune parameters like temperature and max tokens, then chat with your agent in real time to refine its behavior before publishing. ### Visual Workflow Editor The workflow editor lets you build multi-step automations as visual directed graphs. A categorized, color-coded node browser gives you access to over 40 node types, including AI, Logic, Data, Documents, Communication, Sales & Outreach, and more. Drag nodes onto a canvas, connect them, and define how data flows from triggers through AI processing, conditional logic, HTTP calls, and code execution. Workflows can be triggered manually, on a schedule, by webhook, or by platform events. The Sales & Outreach vertical provides purpose-built nodes and templates for B2B lead enrichment, cold email sequences, and CRM integration. ### Real-Time Voice AceTeam agents can talk. Using WebRTC for low-latency audio and Deepgram for transcription, you can have live voice conversations with agents from your browser. Twilio integration enables agents to answer and make phone calls, opening up use cases like customer service lines and appointment scheduling. ### Sovereign Compute Fabric Run AI workloads on your own hardware instead of shared cloud infrastructure. The Fabric connects your GPU servers, databases, and other resources to the AceTeam platform through a secure mesh network. You keep full control of your data while benefiting from centralized orchestration. ### Marketplace Share agents and workflows with other teams, or browse community-built templates. The marketplace accelerates time-to-value by letting you start from proven patterns instead of building from scratch. ### Enterprise Collaboration Organizations get shared workspaces, role-based access, team billing, and audit trails. Invite colleagues, assign permissions, and collaborate on agents and workflows together. ### Projects Projects let you group related agents, workflows, data sources, and other resources under a single umbrella. Each project has its own settings and access controls, making it easy to manage complex initiatives that span multiple assets. See the [Projects guide](/docs/projects) for details. ### Command Center The Command Center is the main dashboard, rebuilt around a jobs-first layout. Active and recent jobs are front and center, giving you immediate visibility into what your agents and workflows are doing right now. ### Conversation Memory Agents can remember past interactions. Conversation Memory automatically indexes prior conversations and makes them available to agents through retrieval-augmented generation (RAG). Returning users get a personalized experience without custom memory logic. ## Who Is AceTeam For? - **Consultants** who want to package AI-powered deliverables for clients - **Operations teams** looking to automate repetitive processes - **Developers** who need a quick way to prototype and deploy AI agents - **Enterprises** that require data sovereignty and on-premise compute ## Architecture at a Glance AceTeam is a web application with a Next.js frontend and a Python backend. The frontend handles the UI, authentication, and API routing. The Python backend handles AI model orchestration, workflow execution, and integration with external services. All AI requests flow through the platform: your browser talks to Next.js, which talks to the Python backend, which talks to AI providers. This architecture keeps API keys secure and enables features like usage tracking, rate limiting, and credit billing. ## Next Steps - [Create your account](/docs/account-setup) and invite your team - [Build your first agent](/docs/first-agent) in the Workbench - [Create your first workflow](/docs/first-workflow) in the visual editor --- ## Account Setup Source: https://aceteam.ai/docs/account-setup Sign up, create your organization, invite team members, and configure billing. Getting started with AceTeam takes a few minutes. This guide walks you through creating your account, setting up your organization, inviting team members, and configuring billing. ## Create Your Account 1. Visit [aceteam.ai](https://aceteam.ai) and click **Sign Up** 2. Authenticate with your email or a supported OAuth provider (Google, GitHub) 3. You will land on the dashboard after signing in New accounts receive **$1 in free credits** to explore the platform immediately. ## Create an Organization Organizations are how AceTeam groups users, agents, workflows, and billing together. 1. Open the organization dropdown in the top navigation bar 2. Click **Create Organization** 3. Enter a name for your organization 4. You are automatically assigned as the organization owner All resources (agents, workflows, compute nodes) belong to an organization, not individual users. This makes it easy to share work across your team. ## Invite Team Members 1. Navigate to **Settings** from the sidebar 2. Open the **Members** tab 3. Click **Invite Member** and enter their email address 4. Choose a role for the invited user Invited users receive an email with a link to join your organization. Once they accept, they can access shared agents, workflows, and compute resources. ## Configure Billing AceTeam uses a credit-based billing system. Credits are consumed when you use AI models, run workflows, and process data. ### Subscription Plans Visit the **Plans** page from the sidebar to view available subscription tiers. Each plan includes a monthly credit allocation and access to different features. ### Manual Top-Ups Need more credits? Purchase them on demand from the **Billing** section in Settings. Your payment method is saved automatically for future purchases. ### Automatic Top-Ups To avoid running out of credits mid-workflow, enable automatic top-ups: 1. Go to **Settings > Billing** 2. Toggle **Auto Top-Up** on 3. Set your **threshold**: the credit balance at which a top-up triggers 4. Set the **top-up amount**: how many credits to add each time When your balance drops below the threshold, AceTeam charges your saved payment method and adds credits automatically. ## Next Steps - [Build your first agent](/docs/first-agent) in the Workbench - [Create your first workflow](/docs/first-workflow) in the visual editor - [Explore the API](/docs/authentication) for programmatic access --- ## Organizations Across Your Devices Source: https://aceteam.ai/docs/organizations-and-devices How your active organization is remembered separately on each device or browser you sign in on. If you belong to more than one organization, AceTeam keeps track of which one is _active_: the workspace whose agents, workflows, and data you're currently looking at. This guide explains how that active organization behaves when you use AceTeam from more than one place. ## Your Active Organization Is Per Device Each device or browser you sign in on remembers its own active organization. Switching organizations on your laptop does not move your phone, and vice versa. That's on purpose. Consultants and agencies live across a lot of client organizations at once. With per-device scoping, your phone can stay parked in Client A while your laptop is heads-down in Client B: no juggling, no accidental cross-posting into the wrong workspace. The unit here is a single sign-in on a device or browser. Multiple tabs of the same browser share one active organization: open three tabs of AceTeam and they all point at the same workspace, which is what you'd expect. ## Switching Is Instant and Local Change your active organization from the organization switcher in the top navigation. The switch is instant and applies only to the device you're on: no re-login, and nothing reloads on your other devices. They keep showing whatever organization you last left them in. ## It Remembers Where You Left Off Your choice is sticky. Close the tab or app and come back later, and that device is right where you left it, still pointed at the same organization. A brand-new device, the first time you sign in somewhere fresh, starts you in your most recently used organization, so you're not dropped into a random workspace. From that point on the new device is independent: switch it wherever you like and it remembers its own choice from then on. You can sign in on as many devices as you want, and each one keeps its own active organization. ## You Only See Organizations You Belong To The switcher only ever lists organizations you're an active member of. If you lose access to the organization a device was showing, say you're removed from a client's org, AceTeam automatically moves that device to an organization you can still use. You won't get stranded looking at a workspace you're no longer part of. ## Pinned Devices Some devices should never drift. A device can be pinned to a single organization, so it always shows that workspace and ignores the per-device switching behavior above. The common case is a wall or lobby display: a screen mounted in an office that's meant to show one organization's board and nothing else. Pinning keeps it locked there no matter who signs in or what happened on anyone's laptop. ## Coming Later Today every device is independent by design. If there's demand, we may add an opt-in that syncs your active organization across all of your devices: switch on one, and the rest follow. For now, each device stands on its own. --- ## Projects Source: https://aceteam.ai/docs/projects Organize agents, workflows, and resources into projects. Projects are the top-level organizational unit in AceTeam. They let you group related agents, workflows, data sources, and other resources under a single umbrella, so everything for a client engagement, product initiative, or internal process lives in one place instead of scattered across the sidebar. ## Creating a Project 1. Click **New Project** from the sidebar or the Command Center. 2. Give the project a name and optional description. 3. Choose the visibility: **Private** (only invited members) or **Organization** (visible to everyone in your org). Projects are scoped to your organization. Each project gets its own page that shows all of its resources at a glance. ## Adding Resources Once a project exists, you can add resources to it: - **Agents**: assign existing agents or create new ones directly within the project. - **Workflows**: link workflows that support the project's goals. - **Data Sources**: attach knowledge bases, files, and document collections the project's agents and workflows can reference. - **Integrations**: connect project-specific API keys, OAuth tokens, or MCP servers. Resources can belong to multiple projects. Removing a resource from a project does not delete it: it just removes the association. ## Project-Level Settings Each project has its own settings panel: - **Default model**: set the default AI model for new agents created in this project. - **Environment variables**: define project-scoped variables that workflows and agents can reference. - **Access control**: manage which team members can view or edit the project and its resources. - **Notifications**: configure alerts for job failures, completed runs, or other project events. ## Managing Projects From the sidebar, click on a project name to open its dashboard. The dashboard shows: - A summary of all resources in the project. - Recent activity: agent conversations, workflow runs, and changes. - Quick actions to create new resources or run workflows. To archive a project, open its settings and select **Archive**. Archived projects are hidden from the sidebar but can be restored at any time. Archiving does not affect the underlying resources: agents and workflows continue to function independently. ## Usage Tracking and Budgets Each project tracks the cost and activity of all its resources in real time. Open a project and switch to the **Usage** tab to see: - **API Cost**: total spend across all agents, workflows, and workbenches in the project. If a budget is set, a progress bar shows how much has been consumed. - **Activity**: counts of agent runs, workflow executions, and workbench sessions. - **Per-User Breakdown**: a table showing each team member's request count and cost contribution. Use the period selector (7 Days / 30 Days / All Time) to change the time window. ### Setting a Budget 1. Open the project's settings. 2. Enter a **Budget Total** (in USD). This is an informational cap: the progress bar will fill as spend approaches the limit. 3. Leave the budget empty for unlimited spend tracking. Budgets are useful for client engagements where you need to monitor costs against a fixed scope. ## Next Steps - [Create your first agent](/docs/first-agent) and add it to a project - [Build a workflow](/docs/first-workflow) within a project context - Learn about [team collaboration](/docs/collaboration) and access controls --- ## Creating Your First Agent Source: https://aceteam.ai/docs/first-agent Build and test your first AI agent in under 5 minutes. # Creating Your First Agent AceTeam makes it straightforward to create a custom AI agent tailored to your specific use case. This guide walks you through building and testing your first agent using the Workbench. ## Open the Workbench From the left sidebar, click **Workbench**. The Workbench is your interactive environment for designing, configuring, and testing agents before deploying them to your team or clients. ## Create a New Agent Click the **New Agent** button in the top-right corner of the Workbench. This opens the agent configuration panel where you define how your agent behaves. ## Choose a Model Select the foundation model that powers your agent. AceTeam supports several leading AI providers: - **OpenAI GPT-4o** -- A strong general-purpose model with broad knowledge and reliable instruction following. Good default choice for most use cases. - **Anthropic Claude** -- Excels at nuanced reasoning, longer documents, and careful analysis. Well-suited for research and writing tasks. - **Google Gemini** -- Offers multimodal capabilities and strong performance across a range of tasks. - **Deepseek** -- A cost-effective option that performs well on technical and coding tasks. Choose based on your requirements. You can always switch models later without losing your configuration. ## Write a System Prompt The system prompt defines your agent's role, personality, and boundaries. This is the most important part of agent configuration. Write a clear description of what the agent should do and how it should behave. For example, if you are building a customer support agent: ``` You are a customer support specialist for a SaaS product. Your role is to help users troubleshoot issues, answer questions about features, and guide them through common workflows. Be concise, friendly, and always suggest next steps. If you cannot resolve an issue, recommend that the user contact the support team directly. ``` Be specific. The more context you provide in the system prompt, the more consistently your agent will perform. ## Adjust Parameters Fine-tune your agent's behavior with these settings in the **Parameters** section: - **Temperature** -- Controls response randomness. Lower values (0.0 to 0.3) produce focused, deterministic answers. Higher values (0.7 to 1.0) produce more creative, varied responses. Start with 0.3 for task-oriented agents and 0.7 for creative ones. - **Max Tokens** -- Sets the maximum length of each response. A typical setting is 1024 tokens for concise replies or 4096 for detailed, long-form output. Leave other parameters at their defaults until you have a reason to change them. ## Test Your Agent Use the chat panel on the right side of the Workbench to interact with your agent in real time. Send a few messages that represent the kinds of questions or tasks your users will bring. Check that responses match your expectations in terms of tone, accuracy, and format. If the responses are not quite right, iterate on your system prompt. Small wording changes can meaningfully shift agent behavior. ## Save and Publish Once you are satisfied with your agent's performance, click **Save** to preserve your configuration. To make the agent available to your team or within workflows, click **Publish**. Published agents appear in the agent library and can be referenced by other parts of the platform. You can return to the Workbench at any time to edit your agent's configuration, test changes, and republish. ## Next Steps Now that your agent is live, consider [creating a workflow](/docs/first-workflow) to automate multi-step processes that use your agent alongside other tools and logic. --- ## Creating Your First Workflow Source: https://aceteam.ai/docs/first-workflow Build a multi-step automation with the visual workflow editor. # Creating Your First Workflow Workflows let you chain multiple steps together into automated pipelines. Using the visual workflow editor, you can build multi-step processes without writing code. This guide walks you through creating and running your first workflow. ## Open the Workflow Editor From the left sidebar, click **Workflows**. This takes you to the workflow dashboard where you can view existing workflows or create new ones. Click the **New Workflow** button to open the visual editor. ## Understanding the Canvas The workflow editor uses a DAG-based (directed acyclic graph) visual canvas. Each step in your workflow is represented as a **node** on the canvas, and **edges** (the lines connecting nodes) define the order of execution. Data flows from left to right through the connected nodes. You can drag to pan around the canvas and scroll to zoom in and out. ## Add a Trigger Node Every workflow starts with a trigger. Click the **Add Node** button or drag from the node palette on the left side of the canvas, then select **Trigger**. This defines what kicks off your workflow. Configure the trigger by clicking on it to open its settings panel. For your first workflow, choose **Manual Trigger** so you can run it on demand while testing. ## Add an AI Node Next, add an **AI** node to the canvas. Click **Add Node** and select **AI**. This node sends a prompt to an AI model and returns the response. Click the AI node to open its configuration panel: - **Model** -- Select the AI model to use (GPT-4o, Claude, Gemini, or Deepseek). - **Prompt** -- Write the instructions for this step. You can reference data from previous nodes using template variables. For example: `Summarize the following input: {{trigger.input}}`. - **Parameters** -- Optionally adjust temperature and max tokens, just like in the Workbench. ## Connect the Nodes Draw an edge from the **Trigger** node's output handle to the **AI** node's input handle. Click and drag from the small circle on the right side of the Trigger node to the small circle on the left side of the AI node. The connecting line confirms that data will flow from the trigger into the AI step. ## Add an Output Node Add an **Output** node to capture the final result of your workflow. Connect the AI node to the Output node the same way you connected the previous nodes. The Output node displays the result when the workflow finishes. ## Run a Test Click the **Run** button in the top-right corner of the editor. Since you used a Manual Trigger, you will be prompted to provide input. Enter a sample value and confirm. The editor highlights each node as it executes. Once the run completes, click on any node to inspect its output in the results panel at the bottom of the screen. Check that the AI node produced the expected response and that the Output node captured it correctly. ## Available Node Types As you build more advanced workflows, you can use additional node types: - **Trigger** -- Starts the workflow (manual, scheduled, or webhook). - **AI** -- Sends a prompt to an AI model and returns the response. - **Conditional** -- Branches the workflow based on a condition, allowing if/else logic. - **HTTP** -- Makes an external API call to fetch or send data to other services. - **Code** -- Runs custom JavaScript for data transformation or business logic. ## Next Steps Save your workflow and experiment with more complex graphs. Try adding a Conditional node to route data based on the AI response, or use an HTTP node to send results to an external service. Workflows can combine any number of these building blocks to automate sophisticated multi-step processes. --- ## Connect AceTeam to Claude, Codex, or Claude Code (MCP) Source: https://aceteam.ai/docs/connect-mcp Connect the AceTeam MCP server to Claude desktop, Codex, or Claude Code on Windows, Mac, or Linux. The connect URL, OAuth sign-in, and fixes when no tools show. AceTeam exposes its agents, workflows, documents, and tools over the Model Context Protocol (MCP). Point any MCP-compatible client at the connect URL below and it can list agents, run workflows, search your knowledge base, and more, all through the same tool interface your client already speaks. This page is the hub: the URL, how to pick your client, and the fixes for the handful of ways a connector can look connected but still show no tools. ## The connect URL ``` https://aceteam.ai/mcp ``` The older path `https://aceteam.ai/api/mcp/aceteam/mcp` still works as a legacy alias, so if you already configured a client with it, it keeps working. ## Pick your client | Client | Windows | Mac | Linux | Recipe | | ------------------- | ------------------------ | --- | ---------------- | ------------------------------------------------------ | | Claude desktop | Yes | Yes | No desktop build | [Connect Claude Desktop](/docs/connect-claude-desktop) | | Codex (CLI and app) | Yes (native or WSL2) | Yes | Yes | [Connect Codex](/docs/connect-codex) | | Claude Code CLI | Yes (PowerShell or WSL2) | Yes | Yes | [Connect Claude Code](/docs/connect-claude-code) | ## Sign in with OAuth Every client authenticates the same way: your browser opens to aceteam.ai, you sign in, and you tick **Allow access across all your organizations** (AceTeam's consent screen; without it, tools are scoped to a single organization). No API key is needed and none is asked for. Never paste an `act_` API key into a connector to skip this step; that key is for headless harnesses, not interactive clients. ## Verify Once connected, ask your client "Who am I on AceTeam?" (this calls the `whoami` tool) or "List my AceTeam agents." A working connection lists well over 100 tools. ## No tools showing? | Symptom | Cause | Fix | | ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | OAuth succeeded, connector says "no tools available" | Connector was saved before the canonical URL went live | Use the URL on this page exactly; remove the connector, fully quit the app, re-add | | "Protected resource does not match expected AceTeam API or Origin" | Token was issued for an old resource identity | Disconnect, reconnect, sign in again (tokens also expire within 7 days anyway) | | Changed the URL but nothing changed | Config is read at launch; `/exit` or closing the window is not a restart | Quit fully and relaunch | | "Authorization failed, check your credentials" toast flashes then disappears | Transient desktop race | Ignore if tools are listed | | Codex: `codex mcp list` shows the server but tools never appear | `codex mcp login aceteam` was never run, or config.toml was edited while the app was open | Run the login, restart Codex | | Claude Code: tools missing in a new session | Server added with project scope in another directory, or stale auth | Re-add with `--scope user`; in `/mcp` clear auth for aceteam and authenticate again | For client-specific detail, see the recipe pages linked in the table above. ## For agents and headless harnesses This page is for a human connecting an interactive client. If you are an AI agent onboarding yourself, or you need an unattended harness that cannot complete a browser OAuth flow, start at [Agent Index](/docs/agent-index) or [Onboarding External Harnesses](/docs/external-harness-onboarding). For a quick copy of the connect URL and config snippets outside the docs sidebar, see [MCP setup](/mcp-setup). --- ## Connect AceTeam to Claude Desktop on Windows and Mac Source: https://aceteam.ai/docs/connect-claude-desktop Add AceTeam as a custom connector in the Claude desktop app on Windows or Mac: Manage connectors, paste the MCP URL, sign in with OAuth, restart, verify tools. ## Before you start You need the Claude desktop app on Windows or Mac (there is no Linux build; Linux users should follow [Connect Claude Code](/docs/connect-claude-code) instead) and an AceTeam account. Custom connectors are available on the Free plan with a one-connector limit, and on Pro, Max, Team, and Enterprise. On Team and Enterprise, an organization owner may need to add the connector under **Organization settings > Connectors** before members can use it. The connect URL: ``` https://aceteam.ai/mcp ``` The older path `https://aceteam.ai/api/mcp/aceteam/mcp` still works as a legacy alias, so an existing connector configured with it keeps working. ## Steps 1. In the chat composer, type `/mcp`. 2. Choose **Manage connectors**. 3. Click **Add custom connector**. 4. Name it **AceTeam**. 5. Paste the connect URL from above. 6. Click **Connect**. [Placeholder: screenshot of the Manage connectors menu] [Placeholder: screenshot of the Add custom connector dialog with the URL pasted] Alternate path, per the Claude help center: **Customize > Connectors > "+" > Add custom connector**. Leave the Advanced settings (OAuth client ID and secret) empty; AceTeam does not need them. ## Sign in (OAuth) A browser window opens to aceteam.ai. Sign in, then tick **Allow access across all your organizations** (AceTeam's consent screen; without it, tools are scoped to a single organization). Click **Approve**. [Placeholder: screenshot of the AceTeam consent screen with the checkbox ticked] No API key is needed. Never paste an `act_` API key into the connector to skip this step. ## Verify Start a new chat, enable the AceTeam connector in the tools menu, and ask "Who am I on AceTeam?" (calls the `whoami` tool) or "List my AceTeam agents." The connector detail should list well over 100 tools. [Placeholder: screenshot of the connector detail showing the tool list] ## Changed the URL? Restart fully `/exit` and closing the chat window are not a restart; Claude keeps running in the background and the connector config is only read at launch. ### Windows Claude keeps running in the system tray after the window closes. Right-click the tray icon and choose Quit, then relaunch the app. [Placeholder: screenshot of the Windows tray icon context menu] ### Mac Open the **Claude** menu in the menu bar and choose **Quit Claude** (Cmd+Q), then relaunch the app. After relaunching, remove and re-add the connector. ## No tools showing? | Symptom | Cause | Fix | | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------- | | OAuth succeeded, connector says "no tools available" | Connector was saved before the canonical URL went live | Use the URL on this page exactly; remove the connector, fully quit the app, re-add | | "Protected resource does not match expected AceTeam API or Origin" | Token was issued for an old resource identity | Disconnect, reconnect, sign in again (tokens also expire within 7 days anyway) | | Changed the URL but nothing changed | Config is read at launch; `/exit` or closing the window is not a restart | Quit fully (see above) and relaunch | | "Authorization failed, check your credentials" toast flashes then disappears | Transient desktop race | Ignore if tools are listed | ## Next - [Connect AceTeam (hub)](/docs/connect-mcp) - [Connect Codex](/docs/connect-codex) - [Connect Claude Code](/docs/connect-claude-code) - [Build your first agent](/docs/first-agent) --- ## Connect AceTeam to Codex (CLI and app) on Windows, Mac, Linux Source: https://aceteam.ai/docs/connect-codex Add the AceTeam MCP server to OpenAI Codex CLI or the Codex app on Windows, Mac, or Linux: one command, OAuth login, config.toml, restart, and troubleshooting. ## Before you start You need either the Codex CLI (`npm install -g @openai/codex`, Node 22) or the Codex app, and an AceTeam account. The connect URL: ``` https://aceteam.ai/mcp ``` The older path `https://aceteam.ai/api/mcp/aceteam/mcp` still works as a legacy alias, so an existing server entry configured with it keeps working. ## Codex CLI 1. Add the server: ```bash codex mcp add aceteam --url https://aceteam.ai/mcp ``` 2. Log in (opens your browser for OAuth; tick **Allow access across all your organizations**): ```bash codex mcp login aceteam ``` 3. Confirm it is connected: ```bash codex mcp list ``` `aceteam` should appear in the list. This writes the equivalent block into `~/.codex/config.toml`: ```toml [mcp_servers.aceteam] url = "https://aceteam.ai/mcp" ``` `auth` defaults to OAuth, so no key or header is needed. `config.toml` is shared by the Codex CLI, the Codex app, and the IDE extension: adding the server once with the CLI makes it available everywhere else that reads the same file. Hand-editing the file does not hot-reload; start a new Codex session or restart the app to pick up the change. ## Codex app 1. Open **Settings > MCP servers**. 2. Click **Add server**. 3. Name it **AceTeam**, choose transport **Streamable HTTP**, and paste the connect URL. 4. Click **Restart**. [Placeholder: screenshot of the MCP servers settings panel] [Placeholder: screenshot of the Add server form with Streamable HTTP and the URL] 5. When prompted, click **Authenticate** and complete the browser sign-in. [Placeholder: screenshot of the Authenticate / Restart prompt] ## Windows notes ### Native (PowerShell) Recommended by OpenAI; the commands above are identical. Config lives at `%USERPROFILE%\.codex\config.toml`. ### WSL2 The same commands work inside the distro. The OAuth callback still lands on `localhost`, and WSL2 forwards it to your Windows browser, so sign-in works the same way. Config lives at `~/.codex/config.toml` inside the distro, which is separate from the native Windows config file. ## Mac and Linux notes Shell differences only; the commands above work as written in bash, zsh, or fish. ## Verify ```bash codex exec "Call the AceTeam whoami tool and print the email" ``` Or ask the same question in an interactive `codex` session. ## No tools showing? | Symptom | Cause | Fix | | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | | `codex mcp list` shows the server but tools never appear | `codex mcp login aceteam` was never run, or config.toml was edited while the app was open | Run the login, then restart Codex | | OAuth succeeded, no tools available | Server entry was added before the canonical URL went live | Use the URL on this page exactly; remove and re-add the server | | "Protected resource does not match expected AceTeam API or Origin" | Token was issued for an old resource identity | Remove the server, re-add, sign in again (tokens also expire within 7 days) | | Changed the URL but nothing changed | config.toml is read at launch | Start a new Codex session or restart the app | ## Headless or agent use For unattended harnesses that cannot complete a browser OAuth flow, see [Onboarding External Harnesses](/docs/external-harness-onboarding) for the `bearer_token_env_var` path. This page stays OAuth-first for interactive human use. ## Next - [Connect AceTeam (hub)](/docs/connect-mcp) - [Connect Claude Desktop](/docs/connect-claude-desktop) - [Connect Claude Code](/docs/connect-claude-code) --- ## Connect AceTeam to Claude Code CLI on Linux, Mac, Windows Source: https://aceteam.ai/docs/connect-claude-code Add the AceTeam MCP server to Claude Code with one command, sign in with OAuth from the /mcp menu, and verify tools on Linux, Mac, or Windows. ## Before you start You need Claude Code installed and an AceTeam account. The connect URL: ``` https://aceteam.ai/mcp ``` The older path `https://aceteam.ai/api/mcp/aceteam/mcp` still works as a legacy alias, so an existing server entry configured with it keeps working. ## Steps 1. Add the server with user scope, so it follows you across projects rather than being tied to the folder you run this in (the default scope is per-project, which is the usual "it worked in one folder only" surprise): ```bash claude mcp add --transport http --scope user aceteam https://aceteam.ai/mcp ``` 2. Start Claude Code: ```bash claude ``` 3. Run `/mcp`, pick **aceteam**, and click **Authenticate** to complete sign-in in your browser. Tick **Allow access across all your organizations** on AceTeam's consent screen, then return to the terminal. 4. Confirm the connection: ```bash claude mcp list ``` `aceteam` should show as connected. ## Verify Ask "Who am I on AceTeam?" (calls the `whoami` tool) or "List my AceTeam agents." ## Changed the URL? ```bash claude mcp remove --scope user aceteam ``` Re-add with the new URL, then start a new `claude` process; the config is read at launch, so an already-running session will not pick up the change. If tools are still missing, open `/mcp`, clear the stored authentication for aceteam, and authenticate again. ## OS notes ### Linux and Mac Identical; the commands above work as written. ### Windows Works the same way in PowerShell or WSL2, with the same commands. ## No tools showing? | Symptom | Cause | Fix | | ------------------------------------------------------------------ | ------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | Tools missing in a new session | Server added with project scope in another directory, or stale auth | Re-add with `--scope user`; in `/mcp` clear auth for aceteam and authenticate again | | OAuth succeeded, no tools available | Server entry was added before the canonical URL went live | Use the URL on this page exactly; remove and re-add the server | | "Protected resource does not match expected AceTeam API or Origin" | Token was issued for an old resource identity | Remove the server, re-add, sign in again (tokens also expire within 7 days) | | Changed the URL but nothing changed | Config is read at launch | Start a new `claude` process | ## Unattended use For headless agents that cannot complete a browser OAuth flow, see [Onboarding External Harnesses](/docs/external-harness-onboarding) and [Agent Index](/docs/agent-index). Humans use OAuth; do not mint an `act_` API key to skip the browser step. ## Next - [Connect AceTeam (hub)](/docs/connect-mcp) - [Connect Claude Desktop](/docs/connect-claude-desktop) - [Connect Codex](/docs/connect-codex) --- ## Ace CLI Setup Source: https://aceteam.ai/docs/ace-cli-setup Install the Ace CLI and run AI workflows locally from your terminal. Ace is a command-line tool for running AceTeam AI workflows on your own machine. It supports 100+ LLM providers through litellm, including OpenAI, Anthropic, Google, and open-source models. ## Prerequisites - **Node.js 18+** for the CLI itself - **Python 3.12+** for workflow execution (the `aceteam-nodes` package) - An API key for at least one LLM provider (OpenAI, Anthropic, etc.) ## Installation Install globally via npm: ```bash npm install -g @aceteam/ace ``` Or run without installing: ```bash npx @aceteam/ace ``` ## Initial Setup Run the setup command to check dependencies and create your configuration: ```bash ace init ``` This will: 1. **Detect Python** -- Verify that Python 3.12+ is available on your system 2. **Install aceteam-nodes** -- Install the Python workflow execution package if not already present 3. **Create config** -- Generate `~/.ace/config.yaml` with default settings ## Configuration The config file at `~/.ace/config.yaml` stores your default settings: ```yaml default_model: gpt-4o-mini ``` Set your LLM provider API key as an environment variable: ```bash # OpenAI export OPENAI_API_KEY=sk-... # Anthropic export ANTHROPIC_API_KEY=sk-ant-... # Google export GOOGLE_API_KEY=... ``` Ace uses litellm under the hood, so any provider supported by litellm works. See the [litellm docs](https://docs.litellm.ai/docs/providers) for the full list. ## Running a Workflow Execute a workflow from a JSON file: ```bash ace workflow run my-workflow.json --input prompt="Explain quantum computing" ``` Options: - `-i, --input ` -- Pass input values to the workflow - `-v, --verbose` -- Show detailed progress messages - `--config ` -- Use a custom config file ## Validating Workflows Check that a workflow file is structurally valid before running it: ```bash ace workflow validate my-workflow.json ``` ## Listing Available Nodes See all node types you can use in workflows: ```bash ace workflow list-nodes ``` This shows the built-in node types including LLM, APICall, TextInput, DataTransform, CSVReader, conditional logic, loops, and comparison operators. ## Example Workflow Here is a minimal workflow that sends a prompt to an LLM: ```json { "nodes": [ { "id": "input", "type": "TextInput", "data": { "text": "Explain AI in one sentence" } }, { "id": "llm", "type": "LLM", "data": { "model": "gpt-4o-mini" } } ], "edges": [ { "source": "input", "sourceHandle": "output", "target": "llm", "targetHandle": "prompt" } ] } ``` Save this as `hello-llm.json` and run: ```bash ace workflow run hello-llm.json ``` ## Troubleshooting - **"Python not found"** -- Ensure Python 3.12+ is installed and available on your PATH. Run `python3 --version` to check. - **"aceteam-nodes not installed"** -- Run `ace init` to install it, or install manually: `pip install aceteam-nodes` - **"API key not set"** -- Export the appropriate environment variable for your LLM provider. ## Links - [Ace CLI on GitHub](https://github.com/aceteam-ai/ace) - [Ace CLI on npm](https://www.npmjs.com/package/@aceteam/ace) - [AceTeam Nodes on GitHub](https://github.com/aceteam-ai/aceteam-nodes) - [AceTeam Nodes on PyPI](https://pypi.org/project/aceteam-nodes/) --- ## Billing & Credits Source: https://aceteam.ai/docs/billing How AceTeam's credit-based billing works. Understand tiers, credits, usage tracking, and how to manage your account. # Billing & Credits AceTeam uses a **credit-based** billing model. All resources (agents, workflows, workspaces, and API keys) are unlimited to create. Usage (running workflows, sending agent messages, making voice calls) consumes credits from your balance. ## Plans | Plan | Monthly Price | Included Credits | Usage Markup | BYOK | | -------------- | ------------- | ---------------- | ------------ | ---- | | **Free** | $0 | Up to $1 (daily) | 20% | No | | **Pro** | $20/mo | $20/mo | 1% | No | | **Max** | $50/mo | $100/mo | 0% | Yes | | **Enterprise** | Custom | Custom | Negotiated | Yes | ### What's included on every plan - Unlimited agents, workflows, and workspaces - Visual workflow builder - Real-time voice interactions - Community marketplace access ### Free plan details Free accounts receive up to **$1 in credits daily**. If your balance drops below $1, we top it up to $1. If your balance is already at or above $1, nothing changes. There is a **lifetime limit of $10** in free credits per account (including the initial signup credit). Once you've received $10 in total free grants, the daily top-up stops. To continue, upgrade to a paid plan or purchase additional credits. ### Paid plan credits Paid plans receive their full credit grant each month when your subscription renews. Credits do not expire but are consumed by usage. ## How Credits Work Every AI operation consumes credits based on the model and amount of work involved. Your available balance is shown in the sidebar and on the [Usage](/usage) page. ### Usage markup When using AceTeam's platform API keys, a small markup is applied based on your plan tier. This covers infrastructure costs (compute, streaming, queue processing). | Plan | Markup | | ---- | ------ | | Free | 20% | | Pro | 1% | | Max | 0% | ### Bring Your Own Keys (BYOK) On AceTeam Max and above plans, you can connect your own API keys for AI providers (OpenAI, Anthropic, Google, etc.). When using your own keys, you pay the provider directly and no platform markup is applied. ## Managing Credits ### Adding credits You can add credits at any time from **Settings > Billing**: - **One-time top-up**: purchase any amount of credits instantly - **Auto-topup**: set a threshold and amount to automatically replenish when your balance runs low - **Upgrade your plan**: higher tiers include more monthly credits and lower markup ### Auto-topup Auto-topup is **off by default**. Turning it on lets you set: - A **threshold** (the balance that triggers a charge) - An **amount** to charge, between **$1 and $500** per charge When your balance drops below the threshold, AceTeam charges your saved payment method for the configured amount. There is a short cooldown between charges so a burst of usage cannot trigger repeated charges back to back. ### Checking your balance Your credit balance is visible in: - The **sidebar** (coin icon with your current balance) - The **[Usage](/usage)** page (detailed breakdown with usage history) ### Grace period If usage pushes your balance negative, AceTeam does not cut you off the instant it crosses $0. Each plan has a grace floor it can dip to before features pause: | Plan | Grace floor | | ---------- | ----------- | | Free | -$10 | | Pro | -$5 | | Max | -$20 | | Enterprise | -$50 | While your balance is negative but still above your plan's grace floor, everything keeps working and the Usage page shows a warning. Once you cross the grace floor, AI features pause until you add credits. You'll see a clear message on the Usage page explaining what happened and how to resolve it. Your agents, workflows, and data are never deleted; they're ready to resume as soon as credits are available. ## Postpaid billing (charge me after) Postpaid billing is opt-in and approval-gated. An organization owner requests it, keeps a valid card on file, and AceTeam approves each organization individually. Until then, nothing changes. Approved organizations can continue working past the normal grace floor, down to their approved exposure cap of $50 to $500. At the end of a billing period, the outstanding deficit is charged automatically to the card on file as a normal Stripe invoice, with a receipt and hosted invoice page. Settlement charges are batched. A deficit below $5 at period close is not charged yet. It rolls forward until the accumulated deficit reaches $5, because Stripe charges 2.9% plus $0.30 per transaction and very small charges are wasteful. Reaching the exposure cap during a billing period triggers an immediate settlement so work can continue. If a settlement charge fails, Stripe retries automatically. Repeated failure pauses postpaid and AI features until the open invoice is paid. Read access to your data is never removed. After payment, access resumes, but postpaid needs re-approval. ## Receipts Every paid Stripe invoice mints a billing receipt: a verifiable record built from your own invoice data (amounts, period, line items) plus an append-only hash chain, so a receipt cannot be silently altered, reordered, or removed without breaking every later receipt in the chain. Receipts are available from **Settings > Billing** and each has a public verification page anyone can use to confirm the receipt's integrity independently. A receipt proves the billing facts match your Stripe invoice. It is **not a tax invoice** and does not itself prove that Stripe processed the charge; that is Stripe's own record. ## Monthly Usage Statement The **Usage** page includes a monthly usage statement: a model-by-model breakdown of what you spent credits on that billing period, downloadable as a PDF. This statement is informational only, a recap to help you understand your spend; it is separate from the verifiable receipts described above. ## Subscription Management Manage your subscription, payment method, and billing history from **Settings > Billing**. You can upgrade, downgrade, or cancel at any time through the billing portal. --- ## Usage Dashboard Source: https://aceteam.ai/docs/usage-dashboard Track every dollar your organization spends across LLM calls, hosted instances, storage, and egress, broken down by container, source, and the user, key, schedule, instance, or webhook that drove the cost. # Usage Dashboard The Usage Dashboard answers a single question: **who used what in which container, and what did it cost?** Every billable resource (LLM gateway calls, hosted instance compute, storage, and egress) rolls up into the logical container that owns it, attributed to the actor that drove the cost. You'll find it under **Settings → Billing → Usage**, or directly at `/organizations//usage`. ## What you'll see The dashboard has three levels: | Level | Page | Purpose | | ---------------- | --------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | | **Organization** | `/organizations//usage` | All containers in the org, ranked by spend, with category breakdown and end-of-month projection. | | **Container** | `/containers//usage` | One container's cost split by source (LLM / Compute / Storage / Egress) and by initiator. | | **Instance** | `/instances//usage` | Resource sparklines (memory, vCPU, egress, volume) plus the LLM calls that originated inside this instance, broken down by model. | Drill down by clicking any container or instance row. Breadcrumbs at the top of each page take you back to the organization overview. ## Reading the numbers ### Cost-plus pricing Every cost figure shown to you is the **billed** number: your organization's tier markup is already applied. Internally we also track the base provider cost so we can keep margin reporting accurate, but you don't need to do the math yourself. | Plan | Markup | | ---------------- | ------ | | Free | 10% | | Pro | 2% | | Max / Enterprise | 0% | ### Smart precision Numbers scale their precision to their magnitude so the dashboard reads naturally: - **$12.40**: 2 decimals for amounts $1 or more - **$0.0420**: 4 decimals between $0.01 and $1 - **$0.000123**: 6 decimals below $0.01 That way a $400 hero total doesn't compete visually with a $0.0003 line item. ### Period selector Use the **Current month** / **Last month** picker in the top-right of any dashboard. The "Updated 2m ago" indicator next to it tells you how stale the loaded data is. Switch the period or hit **Retry** in the error state to refresh. ### Projected end-of-month The "Projected end-of-month" tile is a linear extrapolation from your spend rate so far this period. It's a forecast, not a quote: costs may rise or fall depending on what your agents do for the rest of the month. Hover the info icon on the tile for the same explanation in-app. ## The "Top initiators" view Container detail pages show a **Top initiators** table that mixes five kinds of actors in one ranked list: | Initiator | When you'll see it | | ------------ | ---------------------------------------------------- | | **User** | Interactive run by a logged-in member of the org | | **API key** | Authenticated by one of your organization's API keys | | **Schedule** | A workflow triggered on a schedule | | **Instance** | An LLM call from inside a hosted AceClaws container | | **Webhook** | A workflow triggered by an external webhook | Each row shows the actor's real name (e.g. `prod-bot` instead of `api_key:f3a91b22`) plus a type icon and the billed amount. We resolve names server-side from your organization's own records, so you'll never see another organization's data here. ## Tracked resources ### LLM costs (gateway calls) Captured automatically by the AEP Safety Gateway. Every LLM call your agents, workflows, and instances make through `OPENAI_BASE_URL` is recorded with input/output token counts, model name, and the initiator that triggered it. ### Hosted instance compute For every running hosted instance, AceTeam polls runtime metrics every 5 minutes and writes deltas to the usage table. The four resource types are: | Resource | Unit | What it measures | | -------- | ---------- | ------------------------------------------------------------ | | Memory | `MB·min` | Memory in megabytes integrated over minutes used | | vCPU | `vCPU·min` | Virtual CPUs allocated, multiplied by minutes used | | Egress | `GB` | Outbound network traffic in gigabytes | | Volume | `GB·min` | Persistent volume size in gigabytes, integrated over minutes | Hover the small info icon next to each unit on the instance detail page for the same definition. When you stop or delete an instance, AceTeam runs one final usage flush so the trailing window is captured. You won't lose the last few minutes of activity to the polling gap. ### Storage and egress Tracked the same way as instance compute. Future container types (jobs, batch runs) will plug into the same `instance_usage` table without any UI changes. ## Empty states A brand-new org sees a friendly **"No usage this period"** message instead of a confusing $0.00 screen. Costs will appear as soon as your first agent run, workflow execution, or instance startup is recorded, usually within a few minutes. If a hosted instance hasn't been polled yet (it just started, or the poller hasn't run), the instance detail page shows the same empty state with a hint that samples are recorded every 5 minutes. ## Exporting The container detail page has an **Export CSV** action that downloads a flat list of raw line items for the selected period. Use this for: - Spreadsheet analysis - Importing into your own finance tools - Reconciling with invoices ## What this isn't (yet) The current dashboard is read-only. These features are tracked as follow-up issues: - **Container-scoped cost alerts**: Today you can set org-level alerts. Per-container thresholds are coming. - **Stripe metered billing**: All numbers above are tracked internally; they don't yet flow to Stripe as metered usage. - **Custom date ranges**: Current month and last month only for now. Pre-built periods cover the most common cases. - **Time-series chart**: A spend-over-time chart is on the roadmap. The data is already collected. See the [billing page](/billing) for credits, plans, and BYOK details. --- ## Adoption Ladder Source: https://aceteam.ai/docs/adoption-ladder A behavioural, five-rung readiness ladder for your organization, plus usefulness dimensions beside it, reported by the org_adoption_report MCP tool. # Adoption Ladder Most "is this org adopting AI" questions get answered by self-report: a questionnaire, a survey, a check-in call. The adoption ladder answers it behaviourally instead, from what your organization has already done, read straight from the tables that already record it. No new form to fill out, and no schema change behind it. Call it over MCP with `org_adoption_report`. ## The five rungs Each rung is grounded in a specific table and column, evaluated in order. Your organization's **tier** is the count of consecutive rungs passed, starting from the bottom. | Rung | Passes when | | ---------------------- | ------------------------------------------------------------------------------------------------------------------------ | | **R1 Activated** | At least one completed workflow run or successful agent turn in the window | | **R2 Recurring** | Work happened in at least 3 of the last 4 weeks, including at least one unattended (non-manual) run | | **R3 Side-effecting** | An agent performed at least one side-effecting action (a mutating tool call or a message send), not just reads | | **R4 Multi-actor** | At least two distinct human actors touched the org's work in the window (ran something, approved something, gave feedback) | | **R5 Governed** | At least one human-resolved approval or reviewed workflow step in the window | The report shows every rung PASS up to and including the first FAIL, then a single, specific next step: what to do to climb past that rung. It never lists more than one fix at a time, so there is always one clear next action rather than a checklist. ## Usefulness dimensions Reaching a rung says an org is *doing* the behaviour the rung names. It says nothing about whether the work is any good. Alongside the ladder, the report carries six usefulness dimensions, each with its own sample size (`n`) so a thin sample stays visible instead of being averaged into a false-confident rate: - **U1 Explicit votes**: thumbs up/down on agent chat and other feedback surfaces. - **U2 Run completion**: the share of workflow runs that finished successfully, plus median run duration. - **U3 Re-run**: how often a run gets re-run, split by whether the original succeeded (a want-it-again signal) or errored (a repair loop). - **U4 Approval calibration**: over human-resolved approvals, approve/deny/edit/note rates, queue latency, and **rubber-stamp share** (approved with no edit and no note). A high rubber-stamp share on a "governed" org means the approval gate is present but not doing much adjudicating. - **U5 Feed engagement**: seen and dismissed rates on stored feed cards (insights, suggestions, daily briefs). - **U6 Surface breadth**: how many distinct MCP tools and call types the org actually exercises. None of these dimensions are folded into a single score. A dimension reporting `n=0` means the underlying signal has not landed yet on this deployment, not that the org scored zero. ## Calling `org_adoption_report` | Parameter | Type | Required | Default | Description | | ------------------ | ------- | -------- | ------------ | --------------------------------------------------------------------------------------------------- | | `organization_id` | string | No | | Report a different organization than your own. Requires a platform-admin account. | | `days` | integer | No | `28` | Lookback window, clamped to 7-365. | | `format` | string | No | `"markdown"` | `"markdown"` for the human-readable report, `"json"` for the same data as structured content. | | `scope` | string | No | `"org"` | `"org"` for a single organization's ladder, or `"fleet"` (platform-admin only) for the fleet-wide tier distribution. | The org view never lists user ids, only counts: `R4 Multi-actor` reports a distinct-actor count, never who they are. `scope="fleet"` goes further and never lists organization names either, only a tier distribution and per-dimension medians across every organization with activity in the window. Strictly read-only: this tool never writes, grants, provisions, or notifies. It only reports where an organization sits today. --- ## Publishing Pages and Public Docs Source: https://aceteam.ai/docs/pages-publishing Create and publish a page to a shareable /p/ link, control who can view it (public, unlisted, org, private), manage org pages, attach files, and export to PDF. # Publishing Pages and Public Docs A page is a self-contained HTML document that AceTeam hosts for you at a stable `aceteam.ai/p/{slug}` URL. Use one whenever you hand someone a report, brief, analysis, or any formatted deliverable: instead of pasting long content into a chat, you publish a styled artifact with a shareable link. Pages are created and driven through the AceTeam MCP page tools, so an agent can produce, publish, and revise them without a UI. Every published page is addressed by an unguessable link that carries a high-entropy capability token appended to the slug (`/p/{slug}-{token}`). The slug alone is not a valid address, so a page cannot be found by guessing a memorable slug. ## Creating a page Create a page with `page_create`. You pass a `title` and the page's HTML `content` (a full self-contained document with inline CSS and JS, or a bare body fragment). Instead of inline `content`, you can point the tool at bytes the platform already holds: `content_document_id` (an AceTeam Drive document) or `content_node_id` plus `content_node_path` (a file on your own connected Citadel node). Provide exactly one content source. Useful `page_create` options: - `slug`: the URL slug to use when publishing. Auto-generated if you omit it. - `theme`: a built-in stylesheet (`dark-report`, `light-doc`, or `minimal`) merged into your content so you do not have to re-emit CSS on every page. - `favicon`: an emoji shown as the browser-tab icon. - `description`: a one-line description for sharing previews. - `visibility`: when `publish=true`, sets access to `private`, `org`, `unlisted`, or `public`. Immediate publishing defaults to `unlisted` for compatibility. - `attribution`: defaults to on (a soft "Made with AceTeam.ai" footer). Pass `false` to suppress it on a private or legal artifact. - `chrome`: the toolbar mode for the public render, one of `full` (default), `minimal` (collapsed to a "..." affordance), or `none` (hidden for anonymous viewers). - `tags`: free-form labels for organizing and filtering pages later. A new page is created **private** by default. Pass `publish=true` to create and publish in a single call, which is the common case. You can pass `visibility` with it to choose the access level in the same call. Omit `visibility` when saving a draft. For a page you expect to revise or re-publish under a memorable URL, use `page_upsert` instead. It creates or updates a page by slug and (re)publishes in one call, so re-running with the same slug updates the same URL in place. `page_upsert` also accepts `content_format` (`html`, `markdown`, or `auto`), so you can hand it raw markdown and have it rendered to a styled, sanitized HTML page. ## Visibility levels Visibility is a permission attribute on the page. The token in the URL addresses the page; whether you may actually view it is then decided by the visibility level. The share panel offers a four-rung ladder, in ascending order of exposure. | Level | Share-panel label | Who can view | Link behavior | | ---------- | -------------------- | ------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------- | | `private` | Only you (private) | The creator only. A fellow org member and an anonymous viewer are both denied. | Served `no-store`, never shared-cached. Anyone else is shown a login or "no access" gate, or a 404. | | `org` | Your organization | Any member of the owning organization. | Served `no-store` (varies per viewer). A non-member gets the login or "no access" gate. | | `unlisted` | Anyone with the link | Anyone who holds the token URL, no sign-in required. | Served publicly and shared-cacheable. | | `public` | Public | Anyone who holds the token URL, no sign-in required. | Served publicly and shared-cacheable. | `unlisted` and `public` render identically at serve time. The difference between them is discoverability inside AceTeam's own surfaces (such as the in-org Artifacts index), not who a direct `/p/{slug}` link reaches. On search-engine indexing: because every published page lives behind an unguessable token URL and is not enumerable from its slug, a page is only reachable to a crawler that already has the token link. `private` and `org` pages are gated at serve time and served `no-store`, so a crawler without an authorized session cannot fetch their content at all. An unknown slug and a token-less or wrong-token guess of a real page both return the same 404, so the URL is never an existence oracle for a private page. ### Setting and changing visibility - `page_publish` takes a `page_id` and a `visibility` (default `unlisted`) and returns the public URL. Pass `private` or `org` to keep the page restricted. - `page_upsert` takes a `visibility` argument (default `private`). It never silently downgrades a page you already opened up: the requested visibility only applies while the page is still at the `private` floor, so a page you deliberately made `org` or `public` stays that way across reruns. - `page_unpublish` removes a page from its public URL by flipping it back to `private`. The page and its slug are retained, so a later re-publish restores the same URL. ## Publishing a page: step by step 1. Draft the page with `page_create`, passing the `title` and one content source. Leave it unpublished (the default) if you want to review it first, or pass `publish=true` to publish immediately. 2. Read it back with `page_read` (by `page_id` or `slug`) to confirm the stored HTML, or preview it as its creator at the page's view URL. 3. Publish with `page_publish`, choosing the `visibility` rung you want. This returns the shareable `aceteam.ai/p/{slug}-{token}` URL. 4. Share that URL. A `public` or `unlisted` link opens for anyone; an `org` or `private` link prompts a non-authorized viewer to sign in or shows a "no access" notice. 5. Revise as needed. Use `page_edit` for a small change (exact string replacement against the stored HTML, so copy `old_string` verbatim from a fresh `page_read`), or `page_update` / `page_upsert` to replace the whole document. Each edit creates a new version and republishes in place at the same URL and visibility. Editing never changes who can see a page. Visibility only changes when you call `page_publish`, `page_unpublish`, or pass an explicit promotion through `page_upsert`. ## Finding and managing pages - `page_list` lists your organization's pages, with optional `tags` filtering and `sort` / `order` controls. Tag filtering never widens what a page's visibility already allows you to see. - `page_update` can retag a page without creating a new HTML version: pass only `tags` (pass `[]` to clear all tags). - `copy_page_to_org` clones a page's title and current HTML into a different organization you are an active member of. The copy lands as a new private page there; call `page_publish` in the target org to give it a URL. ## Attaching files Bind an existing Drive file to a page as a downloadable asset with `page_attach_file` (one file) or `page_attach_files` (a batch, capped at 50 ids per call). The file keeps its own storage object; nothing is re-uploaded. Attached files appear in a **Downloads** list in the page's chrome header. A download **inherits the page's visibility**, enforced on every request: a `public` or `unlisted` page yields a link download, an `org` page an org-only download, and a `private` page a creator-only download. Attaching never changes the file's own standalone share link. A page is capped at 50 attachments. To attach a local file, upload it first with `upload_file` (which stores the bytes in your Drive and returns a Document ID), then pass that id to `page_attach_file`. List attachments with `page_list_attachments`, and remove one with `page_detach_file` (the file itself stays in your Drive). ## Exporting to PDF `page_export_pdf` renders a page's current HTML to a PDF, stores it in your Drive, and returns a share link. The link's access **inherits the page's visibility**, so exporting never widens who can see the content: a link-public page yields a link download, an `org` page an org-only download, and a `private` page a creator-only download. A `private` page can be exported only by its creator. Rendering uses the WeasyPrint CSS layout engine, so complex CSS, `@page` headers and footers, and page numbers come through. JavaScript-rendered content is not reproduced, because the export does not run a headless browser. If WeasyPrint's system libraries are unavailable, the export falls back to an HTML-subset renderer where complex CSS does not survive and some non-Latin glyphs may be substituted. Pass `attribution=false` to strip the "Made with AceTeam.ai" footer from a single export without changing the page. ## Ephemeral redirect links `create_ephemeral_page` is a separate tool for a short-lived redirect link, not a content page. Its `content` must be a URL, and it is served at `/go/{slug}` from Redis with a TTL (default 60 minutes). Use `page_create` or `page_upsert` when you want to present an actual document. --- ## Documents, Templates, and E-Signatures Source: https://aceteam.ai/docs/documents-esignature Create a reusable document template, generate a document from it, send it for review, and request an e-signature. Covers merge fields, request signature links, signature status, and retention controls. # Documents, Templates, and E-Signatures AceTeam includes a native document workflow: author a reusable **document template** with merge fields, **generate a document** by filling those fields, optionally **send it for review**, and **request an electronic signature**. The document never leaves AceTeam storage. The signer completes it at an `aceteam.ai/sign/{token}` link, and AceTeam snapshots and hashes the document at send and at completion to produce a tamper-evident record enforceable under US ESIGN/UETA and Ontario ECA. Every step is available as an MCP tool, so an agent can run the whole flow. This page describes the tools and the order they compose in. ## Reusable Document Templates A template is your canonical agreement HTML with `{{field_name}}` merge placeholders in it. Instead of duplicating a contract and hand-swapping every instance-specific value, you fill a handful of named fields. Merge fields are distinct from signing anchors. Anchors like `{{sig}}`, `{{initials}}`, `{{date}}`, `{{sender_sig}}`, and `{{sender_date}}` are left untouched for the signing path to stamp. Keep them in the template exactly as you would in a one-off document. ### Create a template Call `document_template_create` with: - `name`: a human-readable template name, for example "Room lease (3BR)". - `html`: the document HTML, carrying `{{field_name}}` merge placeholders and any signing anchors you want stamped. - `fields`: the field schema, one object per merge field, for example `{"name": "tenant_name", "type": "text", "required": true, "help": "Full legal name", "default": null}`. Field `type` is one of: | Type | Input | Rendered as | | ---------- | ------------------------ | ----------------- | | `text` | any string (the default) | the text as given | | `number` | a plain number | grouped, `1,200` | | `currency` | a plain number | `$1,200.00` | | `date` | strict ISO `YYYY-MM-DD` | `August 1, 2026` | Field names are snake_case. Validation runs entirely at create time and is strict on purpose, so a template can never be authored into a state that produces a broken agreement: - Every `{{placeholder}}` in the HTML must be a declared field or a signing anchor. - Every declared field must appear at least once in the HTML. - An optional field must carry a `default`, because an optional field with no default would render as an empty gap in a signed document. - A declared `default` is type-checked the same way a real value is. A template may declare at most 50 fields, and the HTML is capped at 512 KB. The tool returns the `template_id` and lists the fields it will ask for. ### Inspect, list, and delete templates - `document_template_list` returns a markdown table of `template_id`, name, field count, and creation date. - `document_template_get(template_id, include_html=false)` shows the template's merge fields, which signing anchors are present, and, when `include_html` is `true`, the full HTML. - `document_template_delete(template_id, confirm=false)` soft-deletes a template. The first call previews and changes nothing; call again with `confirm=true` to delete. Signature requests already issued from the template are unaffected, because each one snapshotted its own filled document at issuance. ## Generate a Document and Request a Signature `document_request_signature` snapshots a source document to PDF, computes its SHA-256 (the "sent" hash), persists the request with an unguessable per-signer link, and auto-creates a hosted status Page. Signature requests currently support a single signer. ### Steps 1. **Choose a source.** Pass one of: - `{"type": "template", "template_id": "", "values": {...}}` to instantiate a template and sign the filled copy in one call. `values` maps each declared field name to its value. Substitution is fail-closed: a missing required field, a value that fails its declared type, a value key the template does not declare, or any placeholder still unresolved after substitution fails the call and creates nothing. - `{"type": "page", "page_id": ""}` or `{"type": "page", "slug": ""}` for a hosted AceTeam Page. - `{"type": "document", "document_info_id": ""}` for an uploaded Drive file (PDF or DOCX). Upload it with `upload_file` first so the raw bytes exist. A DOCX is converted to PDF sovereignly on your own Citadel node running the `gotenberg` module, never centrally; if no such node is online the request fails with guidance to install it. - `{"type": "doc", "doc_id": ""}` for an editable `/d/` doc, rendered from its markdown to the same styled artifact the reviewers read. 2. **Name the signer.** `signers` takes exactly one entry: `[{"email": "...", "name": "..."}]`. 3. **Optionally place signature-field boxes.** `fields` is a list of boxes, each `{"kind": "signature"|"date"|"text"|"initials", "page": <1-based int>, "x": , "y": , "width": , "height": , "value": }`. Coordinates are normalized fractions in `0..1` with a bottom-left origin. Omit `fields` to let signing auto-place at a `{{sig}}` anchor, or append a signature page if none exists. Maximum 50 boxes. 4. **Optionally set delivery and counter-signing.** - `message`: a note stored on the signing page. - `expires_in_hours`: a signing-link lifetime; omit for no expiry. - `email_signed_copy_to_signer` (default `false`): email the signer the executed PDF once they complete. - `email_signed_copy_to_requester` (default `true`): email you the same executed PDF on completion. - `cc_emails`: extra addresses CC'd on your copy, up to 10. - `sender`: your own party block on a two-party agreement, for example `{"title": "Director"}`. Your side auto-executes at issuance and is stamped at any `{{sender_sig}}` / `{{sender_date}}` anchor. The name and email come from your authenticated account and cannot be overridden. Pass `{"sign": false}` to record nothing. 5. **Share the link.** The tool returns the `request_id`, the signer's `/sign/{token}` link, the status Page URL, and the sent hash. Creating a request never emails the signer; you share the link yourself. ### What the signer sees The counterparty opens `aceteam.ai/sign/{token}`. The document is served only through the on-domain proxy, never a raw storage URL. Signing is gated behind a one-time code emailed to the signer's registered address, so the signature is bound to control of that inbox. After the signer consents to sign electronically and enters their typed signature, AceTeam assembles an executed PDF (the original, a signature page, and a certificate page enumerating timestamps, IP, user-agent, and both hashes), computes the final hash, and marks the request completed. ## Send a Document for Review First When the parties need to negotiate before anyone signs, share a read-only review link. Reviewing needs only the link, with no email code, because a review link never authorizes signing. 1. **Open a review.** `document_request_review(source, expires_in_hours)` creates an `aceteam.ai/review/doc/{token}` link. The source is one of: - `{"type": "page", "page_id": ""}` or `{"type": "page", "slug": ""}` for a hosted Page. Reviewers select text on the live HTML and leave anchored, threaded comments. - `{"type": "doc", "doc_id": ""}` for an editable `/d/` doc. Its markdown is rendered to HTML and reviewers comment on it the same way. - `{"type": "docx", "document_info_id": ""}` for a stored `.docx` (upload it with `upload_file` first). Its Word tracked changes are imported as the first accept/reject round; resolve them with `review_resolve_change` before promoting. The legacy spelling `{"type": "document", ...}` is still accepted. 2. **Collect feedback.** Nothing binding is created yet. Track reviewer comments with `review_status(review_id)`, reply in-app with `review_reply`, and close threads with `review_resolve_comment`. 3. **Promote to a signature.** When agreed, `document_promote_to_signature(review_id, signers, message, expires_in_hours, fields, email_signed_copy_to_signer)` is your deliberate sign-off. It snapshots the now-agreed document to the immutable hashed PDF and creates the OTP-gated signing link, exactly as `document_request_signature` does. Only the owning organization can promote a review. For a `.docx` review, promotion first requires that every tracked change is resolved and none is an unsupported type, then materializes the clean accepted document and converts that to PDF on your own Citadel node, so the signer signs exactly the accepted changes. ## Check Signature Status `signature_status(request_id)` returns each signer's state (pending, viewed, or signed), the document integrity hashes (sent and final), and, once completed, an on-domain download link for the executed PDF. The download link is a short-lived `aceteam.ai` proxy URL, never a raw storage URL. ## Retention Controls `document_set_retention` sets, clears, or reads a retention policy on a Drive document you own. Retention is a `purge_at` timestamp: at that instant a scheduled sweep auto-deletes the file, marking it deleted and removing it from listings and search. The delete is soft, so an admin can still recover it. Use it to dispose of collected PII, such as ID scans or financial documents, on a schedule. Pick one mode: - **Set by window:** pass `retention_days` (1 to 3650) to purge that many days from now. - **Set an exact instant:** pass `purge_at` as an ISO-8601 timestamp in the future. - **Clear:** pass `clear=true` to keep the file until it is manually removed. - **Read:** pass none of the above to return the current retention unchanged. Retention is org-scoped: you can only touch a document your own organization owns. A `folder_set_retention` tool applies the same policy across every file currently in a Drive folder. ## Rendering HTML to a PDF When you want a document from your own markup rather than a Page or template, `render_pdf(html, css, filename, page_size, margins, assets)` renders styled HTML to a PDF with a real CSS layout engine (WeasyPrint) and stores it in your Drive at organization visibility, returning a scoped share link. Reference images or fonts must be inlined as `data:` URIs or supplied through the `assets` bundle; no external asset is fetched unless the deployment is configured with an asset-host allowlist. Use `page_export_pdf` instead when you want to export an existing AceTeam Page. --- ## Import a Node Folder into Drive and Publish a Public Link Source: https://aceteam.ai/docs/drive-import-publish Pull a folder from your own Citadel node into AceTeam Drive with its subfolder tree intact, then mint one public browse and download link over the result. Files on a connected Citadel node stay on that node until you decide otherwise. When you want the platform to hold a copy, or you want to hand someone outside your organization a link to browse a whole folder tree, you import the folder into your AceTeam Drive and publish a share link over it. This page covers the import and publish primitives (`drive_import`, `folder_share`, `drive_publish_folder`) plus the read helper (`drive_folder_files`) that lists a folder's files for building per-file links, all reachable as MCP tools. ## Import a Folder: `drive_import` `drive_import` reads a folder (or a single `.zip`) from your own connected Citadel node over the secure mesh and recreates it in your Drive. | Parameter | Required | Description | | ------------ | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `node_id` | Yes | A connected Citadel node ID (from `terminal_list_nodes`) to read the folder from. | | `path` | Yes | The folder, or a `.zip`, path on the node, absolute or relative to the node's workspace root. | | `parent_id` | No | Drive folder (`workspace_id`) to recreate the tree under. Omit to import into your Home folder. | | `visibility` | No | Scope for each imported file's own individual share (`link` by default, `org`, or `private`). This is separate from the one link a folder-level share grants; see below. | The subfolder structure is preserved rather than flattened: a source folder with thirty five subfolders lands as thirty five Drive subfolders. Bytes are read from your own node one file at a time, so the import is not bound by a single-call size ceiling, and nothing ever reads the server's own filesystem. The import is idempotent and resumable. A re-run skips a file whose path and byte size are unchanged, without re-reading it. When the size differs, the file is re-read and content-hash checked, and a changed file lands as a new document rather than replacing or dropping anything. One blind spot worth knowing: an in-place edit that happens to preserve the exact byte count reads as unchanged and is not re-imported, since the node's directory listing exposes size, not a modification time. Because of that cheap-skip design, a single call processes up to a fixed number of files. If a large folder hits that cap, the tool tells you to call it again to continue where it left off; a re-run costs almost nothing for the files already imported. ## Mint a Public Link Over an Existing Folder: `folder_share` Once a folder exists in your Drive, `folder_share` mints (or fetches) one public link over it: `aceteam.ai/shared/folder/`. The public page lists every file in the folder and its subfolders, each downloadable and previewable in place. `folder_share` is idempotent on the token: calling it again for the same folder returns the same link rather than creating a second one. A bare call with just `folder_id` is a pure get-or-create; it never downgrades visibility, clears an expiry, or un-revokes a link someone deliberately disabled. Pass a field explicitly to change it: - `visibility`: who may open the link. `link` (the default on a newly minted share) lets anyone with the URL browse; `org` restricts it to members of your organization; `users` restricts it to an explicit `allowed_emails` list; `private` restricts it to you. - `expires_in_hours`: hours until the link expires. A newly minted link is permanent unless you set this. - `disabled`: `true` revokes the link (the public page stops serving) without deleting the folder or its files; `false` re-enables it. Visibility is enforced on the public serving route itself, not a display-only label. The folder ownership check runs before anything is written, so a guessed `folder_id` from another organization is rejected with a generic not-found message rather than confirming it exists. ## The One-Call Composite: `drive_publish_folder` `drive_publish_folder` wraps folder creation, import, and share into a single call: point it at a folder on your connected node and get back a public browse and download link in one step. | Parameter | Required | Default | Description | | ------------ | -------- | ------------------ | --------------------------------------------------------------------------------------------------------------------------------- | | `node_id` | Yes | | Connected Citadel node ID to read the folder from. | | `path` | Yes | | Folder, or `.zip` (copy mode only), path on the node. | | `visibility` | No | `link` | Who may open the link, enforced on the public page: `link`, `org`, `users`, or `private`. | | `name` | No | source folder name | Name for the created top folder. | | `mode` | No | `copy` | `copy` ingests the bytes into central Drive storage; `reference` registers each file by pointer to the node with no bytes copied. | In `copy` mode, the share keeps serving even if the node later goes offline, since the bytes now live in Drive storage. In `reference` mode, the node stays the only place the bytes actually live: the public page streams each file from the node on request, and any file fails to download or preview while the node is offline. This is a deliberate sovereignty tradeoff rather than a bug, and `reference` mode does not accept a `.zip` source, since there is no directory listing to walk inside an unopened zip without ingesting it, which that mode exists to avoid. `drive_publish_folder` is idempotent: re-running it with the same `name` reuses the same top folder and the same share link, skips files that are unchanged, and never duplicates the folder or the link. Like `drive_import`, a very large folder processes up to a fixed number of files per call; re-run to continue, and the link that already exists keeps serving what has landed so far. ## Listing a Shared Folder's Files: `drive_folder_files` Building a per-file "open this PDF" link used to mean scraping the shared folder's public HTML page for each file's id and title. `drive_folder_files` returns the same information as structured JSON instead: every file in the folder and its subfolders, each with its `document_id`, its path relative to the folder, and, when the folder already has a live public share, a ready-to-use direct-open URL. Provide exactly one of: - `folder_id`: a Drive folder your organization owns, listed in owner context. A per-file open URL is included only when the folder already has a live share (mint one with `folder_share` first); this tool never mints a token itself. - `share_token`: a `folder_share` token, listing the files exactly as the public page would, each with its open URL. The listing mirrors the public serving route's scope: only the shared folder's own subtree is listed, never a sibling folder's files, soft-deleted and non-Drive files are excluded, and the walk is depth and size bounded. A `share_token` for a public `link` folder is listable by anyone; for a scoped visibility (`org`, `users`, `private`) the caller's organization must own the folder, which is deliberately stricter than the serving route's per-email `users` grant, so a leaked scoped token cannot be used to enumerate another organization's files. A disabled or expired share, an unknown token, or a folder in another organization all read as an identical generic not-found result. ## Typical Flow 1. Connect a Citadel node and confirm it is online with `terminal_list_nodes`. 2. Call `drive_publish_folder` with the node id and the folder path to import the tree and mint one link in a single call, or call `drive_import` followed by `folder_share` if you need to control the import and the share separately. 3. Call `drive_folder_files` with the resulting `folder_id` or `share_token` to build per-file open links, instead of parsing the public page's HTML. 4. If the import reports it hit its per-call file cap, re-run the same call to continue; unchanged files are skipped and the existing link keeps serving in the meantime. --- ## Manage Google Calendar Events and Bookable Scheduling Links Source: https://aceteam.ai/docs/calendar-booking Read, create, update, and search Google Calendar events, find free time, and (once enabled) publish a Calendly style booking link backed by your own calendar. Calendar access has two layers. The first, connecting a Google account and working with its events, is live today. The second, a bookable scheduling link an invitee can use to grab a slot on your calendar without an AceTeam account, is authored end to end but waiting on a database migration before it goes live. This page covers both, and is explicit about which is which. ## Connect a Google Account Call `google_connect` to get an OAuth URL and connect a Google account for calendar (and Drive) access. `google_list_accounts` lists every connected account with its email, label, and default status, and `google_disconnect` revokes access for all connected Google accounts. Every calendar tool below accepts an optional `account` parameter (an email or label) to target a specific connected account; omit it to use the default one. ## Read and Search Events - `calendar_list_events` lists events in a time range (`time_min`, `time_max`, both ISO 8601, required), with `max_results` (default 10, max 250) and `calendar_id` (default `primary`). - `calendar_search_events` searches events by text `query`, optionally bounded by `time_min`/`time_max`. - `calendar_suggest_times` finds open slots by checking free and busy data, given a `duration_minutes` (default 30) and an optional search window (defaults to now through three days out). ## Create, Update, and Delete Events - `calendar_create_event` takes a `summary`, `start_time`, and `end_time` (ISO 8601), plus optional `description`, `location`, `attendees` (a list of emails), and `timezone_str` (default UTC). - `calendar_update_event` takes an `event_id` and changes only the fields you pass; everything else is left as is. - `calendar_delete_event` removes an event by `event_id`. All three take the same optional `calendar_id` (default `primary`) and `account` parameters as the read tools. ## Bookable Scheduling Links (Not Yet Enabled) A booking type is a reusable, Calendly style scheduling link: you configure a duration, availability windows, and a backing calendar once, and share the link so someone can pick an open slot without ever signing in to AceTeam. The MCP tools that manage booking types, `booking_type_create`, `booking_type_list`, `booking_type_get`, `booking_type_update`, `booking_type_delete`, `booking_temp_link_create`, and `booking_list`, are built and registered, and a public invitee facing booking page already exists in the web app to serve a minted link. However, the Supabase tables that back these tools are authored but not yet applied, pending schema approval. Until that migration lands, every one of these tools returns a plain "Booking storage is not set up yet" message rather than doing anything. Treat this section as documentation of a capability that is wired but switched off, not a feature you can use today. Once enabled, the shape will work like this: ### `booking_type_create` Creates a bookable link, scoped to your organization and owned by you. Key fields: - `handle`: the public link slug, unique within your organization (lowercase letters, digits, and hyphens). - `title`, `duration_minutes`, `timezone`: what invitees see and how long the meeting runs. - `availability`: weekly windows plus date overrides, expressed as `{"weekly": [{"weekday": 1, "start": "09:00", "end": "17:00"}], "overrides": [{"date": "2026-09-01", "windows": []}]}`, where weekday runs 0 for Sunday through 6 for Saturday and an override with an empty `windows` list closes that date entirely. - `buffer_before_minutes`, `buffer_after_minutes`, `min_notice_minutes`, `max_per_day`: padding and rate limits around bookings. - `medium` and `medium_detail`: `auto_video`, `phone`, `in_person`, or `custom`, plus free text detail for it. - `invitee_questions`: a list of questions asked at booking time, each with a `key`, `label`, `type`, and `required` flag. - `calendar_provider`: `google`, `microsoft`, or `internal`. Google and Microsoft route through a connected account and send a real calendar invite to the invitee. The internal calendar does not email invitees at all; a booking against it is display only for them. - `calendar_account`: the connected account (email or label) that a `google` or `microsoft` provider routes through. Configuring a real Google or Microsoft account that is not actually connected fails at creation time rather than silently falling back to the internal calendar. `booking_type_list`, `booking_type_get`, `booking_type_update`, and `booking_type_delete` do the expected list, read, partial update, and soft delete against your organization's booking types. A handle freed by deleting a type becomes available for reuse. ### `booking_temp_link_create` Mints a temporary, single-use or time-limited link bound to one booking type, for sending to a specific invitee without exposing the type's stable public handle. It accepts a `mode` of `single_use` (stops working after one booking) or `expiring` (reusable until a set expiry), an `expires_in_minutes`, and an optional `slot_shortlist` of specific ISO 8601 start times to offer instead of the type's full availability; each shortlisted slot is still re-checked against real availability at booking time. ### `booking_list` Lists who booked what, filterable by booking type, status (`booked`, `rescheduled`, or `cancelled`), and a date range. Times are handled in UTC throughout. --- ## Parse PDFs and Scanned Documents with Sovereign OCR Source: https://aceteam.ai/docs/ocr-document-parsing Turn an uploaded PDF or image into structured text on your own fabric, one file or a whole directory at a time, with no central OCR fallback. `file_parse` and `file_parse_batch` extract structured text from a PDF or image. Both run entirely on your own fabric: a born-digital PDF page (one with real embedded text) is extracted on CPU, and a scanned page or image is routed to a vision-language OCR model on one of your Citadel GPU nodes. There is no central cloud OCR fallback. ## Parsing One File: `file_parse` Provide exactly one input source: | Parameter | Description | | ----------------------- | --------------------------------------------- | | `document_info_id` | A file already in your AceTeam Drive | | `url` | A signed aceteam.ai share URL | | `node_id` + `node_path` | A file on one of your connected Citadel nodes | Other parameters: `model` (defaults to `baidu/Unlimited-OCR`), `max_pages` (default 5, the number of PDF pages processed in one call), `text_yield_threshold` (default 50, the per-page character count below which a page is treated as a scan and routed to OCR instead of trusted as extracted text), `include_bboxes` (default true, per-block bounding boxes in the result), and `retention_days` (default 30, the auto-delete window for a Drive input; pass `0` to keep the file). Routing is per page, not per document: within one PDF, a page with real embedded text extracts on CPU while a scanned page in the same file routes to OCR. ## Parsing a Directory: `file_parse_batch` `file_parse_batch` is the batch form for a folder of scans on a Citadel node rather than one file per call. It lists the directory once, resolves and pins a single node to serve the whole batch (so a batch of many files triggers at most one model warm-up wait, not one per file), and reuses the same parsing core `file_parse` uses for each file. The call waits inline for up to about 25 seconds. A small batch that finishes in that window returns its full result with `status: "done"`. A batch still running returns a `job_id` and `status: "running"` instead; poll it with `file_parse_batch_status(job_id)`, or call `file_parse_batch` again with the same `job_id` to both continue the batch and get the latest status in one call. One bad file never aborts the batch: a page that fails to parse is recorded as failed and the rest of the directory still gets processed, and a file that failed on an earlier pass is retried on resume. A systemic failure, such as the OCR model still warming up, no online node, or the pinned node dropping mid-batch, stops the current pass immediately instead of burning the whole per-call file budget re-discovering the same failure file after file; resume once the underlying problem clears. By default the batch result is a compact manifest (status, engine, page count, character count per file), not the full parsed text, to avoid an oversized response on a large batch. Pass `out_dir` to also have the full parsed JSON for each file written back to the node itself. ### Starting or Resuming a Batch | Parameter | Description | | ----------------------- | ------------------------------------------------------------------------- | | `node_id` + `node_path` | Required to start a new batch | | `job_id` | Resume a prior batch; every other argument is ignored | | `pattern` | Optional glob filter on filenames, e.g. `*.png` | | `confirm` | Required to dispatch a batch the cost preview gate holds for confirmation | ### Cost Preview Starting a new batch that is both non-trivial (more than 10 files) and non-free (the target node is shared with your org rather than owned by it; an own-node batch always settles at zero cost and skips this) returns a cost estimate and preview instead of dispatching anything. Call again with `confirm=true`, either on the first call or with the returned `job_id`, to run it. ### Polling: `file_parse_batch_status` `file_parse_batch_status(job_id)` is a read-only poll that does no parsing itself. If the job is not actively being processed (`status: "paused"`), polling alone will not make progress; call `file_parse_batch(job_id=...)` again to resume it. Pass `format="summary"` to get only the counts (total, done, failed, and similar) instead of the full per-file manifest, which is the cheaper shape for watching a long batch. --- ## Store Structured Agent Records Without a Custom Database Source: https://aceteam.ai/docs/records Write, query, and list small JSON records under a namespace and key, with secret scrubbing on write and scoped access to another user's data. Not every piece of state an agent needs is a document, a memory, or a workflow run. Sometimes it is a small structured fact: a processed order id, a computed score, a pointer to something else. Records give agents a lightweight, namespaced place to keep that kind of data without standing up a real database, a schema, or joins. A record is one JSON value stored under a `namespace` and a `record_key`, optionally grouped under a `group_id`. Writing the same `namespace`, `group_id`, and `record_key` again updates the record in place instead of creating a duplicate, which is what makes records safe to write repeatedly from a retryable job. ## Write a Record: `record_write` | Parameter | Required | Default | Description | | ----------------------------- | -------- | -------- | --------------------------------------------------------------------------------------------------------- | | `namespace` | Yes | | Logical family for the record, for example `memory.note` or `trainer.config`. | | `record_key` | Yes | | The idempotency key within the namespace and group. | | `value` | No | `{}` | The JSON body of the record. | | `kind` | No | `record` | Leave at the default; other backend kinds require locators this tool does not expose. | | `group_id` | No | | Parent or grouping id. Omit for a top level record. | | `scope` | No | | Optional sub-scope folded into the grouping. | | `seq`, `label`, `flag`, `num` | No | | Optional typed scalar fields, whose meaning is up to your namespace, useful as filters on `record_query`. | | `source` | No | | Optional free text label recorded on the row. | Ownership is never something you pass in. The organization, the owning user, and the client (always `mcp` for this surface) are taken from your authenticated session, so a record you write can never be attributed to a different user or a different organization. Before anything is stored, the `value` is passed through the platform's secret scrubber, which redacts things like API keys, `.env` style secrets, and high entropy tokens. If scrubbing itself cannot complete, the write fails closed: nothing is stored, rather than risking an unscrubbed secret landing in a record. ## Read and List Records - `record_read` reads one record by its id. - `record_query` is a filtered, newest first search over your records, matching on `namespace`, `kind`, and the typed scalar fields (`label`, `flag`, `num`, `seq`). Page through results with `cursor`, set to the previous page's last `created_at`. - `record_list` is the cheaper "show me this owner's records" surface: filter by `namespace` and `group_id` with no scalar filtering, for when you already know roughly what you are looking for. All three default to your own records. Reading or listing another user's records requires passing `subject` (their user id) explicitly, which in turn requires your API key to carry the `memory:read_cross_user` scope. That check is enforced against the record actually fetched from the database, not against whatever `subject` argument you supplied, so a bare record id can never be used to read across users by accident. ## Delete a Record: `record_delete` Deletes are soft: the record is tombstoned rather than physically removed, which keeps history and observability intact. Deleting another user's record follows the same `memory:read_cross_user` rule that reading one does. ## Scopes Reading (`record_read`, `record_query`, `record_list`) requires the `memory:read` scope; writing and deleting (`record_write`, `record_delete`) require `memory:write`. Both are standard API key scopes you set when minting a key with `create_api_key`. ## When to Use Records Versus Other Storage Records are the right fit for a small, structured, per-agent or per-user fact that needs to be looked up later by key, without the overhead of a knowledge collection or a Drive file. For searchable free text, use a knowledge collection with `upload_document`, covered in [Ground Agents with Knowledge Collections (RAG)](/docs/knowledge-collections). For raw file bytes, use `upload_file`. For a durable markdown note meant to persist across sessions and be recalled by an agent, use the memory tools (`memory_write`, `memory_search`) described in [Agent Memory](/docs/agent-memory). --- ## Publish a Hosted HTML App at a Public Short Link Source: https://aceteam.ai/docs/hosted-apps Create, update in place, and delete a self-contained HTML app served at a public short code, with the same visibility rules as hosted pages. An Ace App is a self-contained HTML application served publicly at `/a/{short_code}`, ships the `window.aceApp` SDK for reading and writing structured data, and runs with no content security policy sandbox. It is the same hosted app surface the web app builder produces; `create_app` and its siblings give an agent that same capability directly, so the resulting app is indistinguishable from one built through the UI. ## Create an App: `create_app` | Parameter | Required | Default | Description | | --------------------------------- | ------------ | --------------- | ------------------------------------------------------------------------------------- | | `title` | Yes | | App title, also the name of its folder in your organization's Ace Apps. | | `html` | One of three | | The app's HTML, given inline. Stored as `index.html` and served at the root. | | `html_document_id` | One of three | | An AceTeam Drive document id whose bytes become the app's HTML, resolved server side. | | `html_node_id` + `html_node_path` | One of three | | A file on a connected Citadel node to read the HTML from directly. | | `visibility` | No | `unlisted` | `private`, `org`, `unlisted`, or `public`. | | `target_organization_id` | No | your active org | Create the app in a different organization you belong to. | Provide exactly one content source: `html` inline, a Drive document id, or a node file. This means HTML the platform already holds, a file a node just wrote, a Drive document, never has to be retyped as a tool argument. `visibility` gates the public `/a/{short_code}` link itself, matching how hosted pages already work: `private` serves only to you, its creator; `org` serves only to active members of the owning organization; `unlisted` and `public` both stay openly reachable to anyone who has the code, and differ only in whether the app shows up in the organization's own app index (`private` and `unlisted` show it only to you and organization admins; `org` and `public` show it to every member). ## Update an App in Place: `app_update` `create_app` mints a new short code on every call, which orphans the old URL and the views and engagement it had accrued. `app_update` instead appends a new version to the existing app and repoints its short code to that version, so the public URL, and every analytics row keyed on it, is preserved. It mirrors how republishing a hosted page at the same slug works. Only `index.html` is replaced. Text sibling files from the current version, such as a `style.css` or `app.js`, carry forward unchanged, so a small HTML tweak never drops them. A binary asset, such as an image, in the current version is not carried forward; if the app depends on a binary sibling file, recreate it with `create_app` instead. The app is resolved by `short_code` within your organization only. An unknown code and one owned by another organization read identically, so this can never be used as a way to probe whether a code exists elsewhere. ## Delete an App: `app_delete` The cleanup counterpart `create_app` never had: an app made to test, demo, or iterate on can be removed instead of sitting permanently in the organization's Ace Apps folder. Deletion follows the platform's standard two-step confirmation. The first call, with `confirm` left at its default of `false`, returns a preview and changes nothing. Call again with `confirm=true` to actually delete. The delete itself is soft: it sets a deletion timestamp rather than dropping any row, so the app and its files stay available for audit or restore, while the public `/a/{short_code}` route stops serving immediately, since it gates on that same column. Deleting an already deleted app is a clean no-op that reports `already_deleted` rather than an error. ## A Ready-Made Template: `create_survey_app` `create_survey_app` builds a fixed, ready-to-clone student AI-usage survey and starts collecting responses in one call, on top of the same `create_app` path. It takes a `title` (the survey heading, typically a class name), an optional `course_name` or custom `intro`, a `collection` name for where responses land (default `responses`), and the same `visibility` and `target_organization_id` options `create_app` takes. The survey is anonymous: a visitor needs no AceTeam account, and each submission appends one row. Response collection is deny by default, so an app on its own collects nothing; `create_survey_app` turns it on for the new survey's collection automatically, the same write `hosted_surface_config` performs directly. Only your organization can read results back, through the owner responses view, a CSV export, and the `/surface-data` page. ## Visibility Model Hosted apps and hosted pages share one visibility model: | Visibility | Who can view | | ---------- | --------------------------------------- | | `private` | The creator only | | `org` | Members of the owning organization only | | `unlisted` | Anyone with the link | | `public` | Anyone with the link | `private` and `org` are enforced at serve time on every request, not just hidden from a listing: a request for a private or org-scoped app resolves the viewer's session and checks it against the app's owner and organization before any bytes are returned. A denied request for a private or org app reads identically to a bad short code, since apps carry no separate capability token to distinguish a denied viewer from one asking about a code that never existed. --- ## Route Chat by Team and Post to Channels with Team Chat Source: https://aceteam.ai/docs/teams-team-chat Group members into teams for chat routing, then create channels, open conversations, and post messages on AceTeam's native Team Chat surface. Teams and Team Chat are two related but distinct constructs. A team is a membership and routing group: a named set of humans and agents that can be attached to a channel so everyone on the team can reach it. Team Chat is AceTeam's native chat surface itself, the channels, conversations, and messages the `/chat` UI renders. You can post to a channel without ever creating a team; teams exist to make "everyone on this team can see this channel" a one-step grant instead of adding members one by one. ## Teams: Membership and Routing A team has a name, a handle, an accent color, a list of members (each tagged `human` or `agent`), and a list of channels it is attached to. - `list_teams` lists every team in your organization with its handle, accent, member count, and channel count. - `get_team` returns one team's full detail: members and channels. - `add_team_member` adds a human or agent to a team. Requires org admin, and the member must belong to your organization. - `remove_team_member` removes a member from a team. Requires org admin. Attaching a team to a channel is done from the channel side (see below), not from the team tools. ## Team Chat: Channels, Conversations, and Messages A channel is a container for a list of conversations, and each conversation is a self-contained message tree. Reading a channel by default spans every conversation inside it, each message labeled with which conversation it belongs to. ### Who Can Read and Post Access to a channel follows one rule, consistently across every tool: - An **effective member** can read and post. You are an effective member of a channel if you have a direct roster row on it, or if you belong to a team that is attached to the channel. - A channel marked visible to non-members is **readable** by any organization member, but not postable, if you are not an effective member. - An **archived** channel is not reachable at all, by anyone. ### Listing and Reading - `team_chat_list_channels` lists every channel you can open: every channel of your org visible to non-members, plus any channel you belong to directly or through a team. Each row shows name, ID, visibility (`organization`, `private`, or `direct`), and description. - `team_chat_read_messages` reads recent messages from a channel, oldest first, up to 100 at a time. Pass `since`/`until` (ISO 8601 or a relative duration like `24h`) to bound the window, `before` to page backwards from a message ID, or `conversation_id` to read one conversation instead of the whole channel. - `team_chat_search_messages` runs a case-insensitive substring search over message text across every channel you can open, optionally restricted to one `channel_id`. - `team_chat_list_members` lists a channel's effective roster (direct members plus anyone reachable through an attached team, each tagged with its source), or every active member of your organization when `channel_id` is omitted. ### Posting `team_chat_send_message` posts as the calling user and requires channel membership; a read-only channel rejects the post. The message goes to one of three places depending on what you pass: - Neither `parent_message_id` nor `conversation_id`: opens a **new conversation**, titled from the first line of the message body. - `conversation_id`: appends to an **existing conversation**, returned by the read and search tools. - `parent_message_id`: **replies** to a specific message inside its conversation. The response always carries both the message ID and the conversation ID, so you can follow up with either form on a later call. ### Creating a Channel `team_chat_create_channel` uses the same creation path as the web UI: you become the channel admin, `initial_member_ids` are added as members, and the org's Ace agent is auto-added so it is mentionable. Any organization member may create a channel. Visibility is `organization` (readable org-wide) or `private` (members only). There is no name-uniqueness constraint, so two channels can share a name; the returned ID is what disambiguates them. ### Archiving a Channel `team_chat_archive_channel` is the paired reversal of channel creation, and the same operation the UI's delete performs: it stamps the channel as archived so it leaves listings while its message history is preserved. **There is no un-archive path.** It requires you to be an admin of the channel (a direct `admin` roster row, which the creator always has, or an org owner/admin), and uses two-step confirmation: the first call (`confirm=false`, the default) previews and changes nothing; call again with `confirm=true` to archive. ## Typical Flow 1. `team_chat_create_channel` to make the channel, naming any humans or agents who should start as members. 2. Optionally build a `team` with `add_team_member` and attach members to it, so future channel access can be granted by team rather than one member at a time. 3. `team_chat_send_message` with no `conversation_id`/`parent_message_id` to open the first conversation. 4. `team_chat_read_messages` or `team_chat_search_messages` to catch up on what's happened, and `team_chat_send_message` with `conversation_id` or `parent_message_id` to continue or reply. --- ## Keep a Shared Daily Journal for Humans and Agents Source: https://aceteam.ai/docs/journal Append dated, sectioned entries to a per-user journal from an agent or a human, read them back in one canonical format, search across days, and share a day or a single section. The journal is a per-user, date-keyed daily log. Both humans and agents write to it, appending sections through the day, and it is designed to be read back as a running record rather than a single flat document. Entries live in first-class tables (a journal entry per day, sections within it, and shares), not in the knowledge base, so reads are structured and cheap. ## How Entries Are Structured Each day has one entry. Within that entry, sections accumulate in order as they are appended. Every section carries an occurrence timestamp, a heading, tags, and the markdown content itself. **Every MCP write is attributed to an agent.** The caller cannot mark a section as human-authored; the write path always stamps the section with the calling client's identity as an agent author. This is a deliberate provenance rule, not an oversight, so a reader always knows whether a given section was typed by a person in the web app or appended by a tool call. ## Appending to the Journal `journal_append` resolves or creates the day's entry (default: today) and appends a new section at the end. | Parameter | Required | Default | Description | | ------------- | -------- | ------- | ------------------------------------------------------------------- | | `content` | Yes | | Markdown body of the section | | `heading` | No | derived | Short title; auto-derived from the first line if left blank | | `date` | No | today | `YYYY-MM-DD`, so you can file late-night work under the right day | | `occurred_at` | No | now | ISO 8601 timestamp for when it happened, backdatable within the day | | `tags` | No | none | Tag strings | | `timezone` | No | `UTC` | IANA timezone used to resolve "today" | | `source` | No | `mcp` | Client identity label recorded as the author, e.g. `claude-code` | Appending the exact same content to the same day twice is a no-op: the write is deduplicated by a content hash, so a retried call cannot double up a section. ## Reading the Journal `journal_read` returns entries in one canonical markdown format: an H2 heading per day, and a `### HH:MM [kind: label] heading` line per section, with each section's UUID embedded as an HTML comment for later reference. This is deliberately a single, stable shape so that reading a range of days concatenates into one fetch rather than a list of separately-shaped documents. Call it with no arguments to get today's own entry, or pass `date` for one specific day, or `start_date`/`end_date` for a range (capped at 31 days). Pass `user` (a teammate's email or user ID) to read their org-visible entries and sections instead of your own. For a cheap index without pulling every section's body, use `journal_list` first: it returns just the headings for a window of days (default the last 30 through today), which is the list-then-read pattern for scanning what exists before fetching the content. `journal_search` runs a case-insensitive keyword search over section headings and content, optionally bounded by `start_date`/`end_date`, and optionally including org-visible sections of teammates with `include_org=true`. ## Sharing a Day or a Section `journal_share` creates, or fetches if one already exists, a shareable link for a whole day or a single section, returning a `/j/{token}` URL. - Pass `date` (defaults to today) to share the whole day, or `section_id` to share just one section. - `visibility` is `link` (anyone with the URL) or `org` (organization members only). - `expires_in_hours` optionally bounds the link's lifetime. The call is idempotent per target: sharing the same day or section again returns the existing link rather than minting a duplicate. ## Typical Flow 1. `journal_append` through the day as work happens, letting the heading auto-derive or setting one explicitly. 2. `journal_list` to see what days have entries, then `journal_read` for the days or range you want to review. 3. `journal_search` when you remember a keyword but not the date. 4. `journal_share` to hand a day or a section to someone outside the loop, scoped to `link` or `org` visibility. --- ## Push Curated Cards to the Home Feed and Manage What Surfaces Source: https://aceteam.ai/docs/feed Surface a report or a suggestion on your reading-list feed, opt a hosted page out of the feed permanently, and read the same assembled feed the web and iOS apps render. The feed is your per-user reading-list home stream: the surface that mixes activity, pending approvals, and curated cards into one ranked list. Publishing a hosted page auto-mints a feed card, private to its author until they promote it, so most of what shows up happens without a separate call. These tools let an agent curate the feed directly, for the cases where a page needs promoting or something needs to be pushed proactively. ## Pushing a Card `feed_push` puts a curated card on the feed: a headline, an optional "why it matters" line, and an optional longer body. A card points at exactly one destination: - `page_id`, to link the card to one of your organization's published hosted pages, or - `url`, an external `http(s)` destination when there is no page. These are mutually exclusive. `kind` is `insight` or `suggestion` (`daily_brief` is reserved for the platform's own digest service). `audience` is `me` (private, the default) or `org` (every member); an `org` audience is rejected if the underlying page is not shared to the organization. `expires_in_hours` optionally ages the card off the feed automatically. A card whose dedupe key already exists is a no-op, so re-pushing the same headline and destination will not duplicate it. ## Opting a Page Out of the Feed Some published pages, a signed legal document, a one-off form response, should never surface as a feed card even though publishing normally auto-mints one. `feed_suppress` marks a page as suppressed and removes any of its existing feed cards. The suppression survives republishing the page, so a signable or legal page never quietly reappears in someone's feed after an edit. Pass `suppressed=false` to clear the flag; clearing it does not retroactively re-mint any card that was removed. ## Reading Your Own Feed `feed_read` is strictly read-only and always reads the calling user's own feed: organization and user identity come from the authenticated MCP context, never from an argument, so there is no way to read someone else's feed through this tool. By default it returns the fully assembled feed: activity, pending approvals, and stored curated cards together, ranked and paginated the same way the web and iOS feed actually render them, since it is fetched through the same internal path those clients use rather than a separate reimplementation. The result is a compact markdown table plus a summary of counts by `kind` and by source kind, and the stored unread total. If the live, fully-assembled feed cannot be reached (a misconfigured deployment or a network failure), the tool degrades to the stored half only, meaning `insight`, `suggestion`, and `daily_brief` cards, with a loud in-band warning ahead of the table. It never silently reports the stored half as if it were the whole feed. Parameters: `limit` (default 30, clamped to 1 through 100), `unseen_only` (only cards you have not marked seen; applies to stored cards only, and cannot be combined with `kind="activity"` or `kind="approval"`), and `kind` to filter to one of `activity`, `approval`, `insight`, `suggestion`, or `daily_brief`. ## Typical Flow 1. Publish a hosted page normally; a private feed card is minted for you automatically. 2. `feed_push` when you want to surface something that has no page of its own, or to promote a page's card to the whole organization. 3. `feed_suppress` on any page (a signed contract, a one-time form) that should never appear as a feed card, including after future republishes. 4. `feed_read` to self-diagnose exactly what is showing up in your own feed, and why, if something looks missing or unexpected. --- ## Agentic Execution Protocol Source: https://aceteam.ai/docs/aep-overview The protocol layer that makes multi-organization AI workflows accountable. Start here to understand why AEP exists and how to read the full specification. # AEP: Agentic Execution Protocol AI agents are calling other AI agents across organizational boundaries. A consulting firm's strategy agent calls a research firm's analysis agent, which calls an analytics firm's data processing agent. Three organizations, three runtimes, one workflow -- and no shared accountability layer. In human professional services, this problem was solved decades ago with paperwork: receipts track costs, footnotes trace conclusions to sources, and NDAs control who sees what. AEP brings that same accountability to AI agent workflows as a vendor-neutral wire protocol. ## Why This Matters If you're building, buying, or evaluating AI agent systems, you will encounter this problem. The moment agents cross organizational boundaries -- and they will -- you need answers to three questions: 1. **Who spent what?** When a workflow spans three organizations, who pays whom? How do you settle costs when each participant has their own billing system? 2. **Where did this conclusion come from?** When an AI agent says "the drug pipeline is promising," can you trace that claim back through intermediary analyses to the original data? Regulators are already asking this question. 3. **Who was allowed to see what?** Patient data marked as protected health information shouldn't reach an agent whose organization hasn't signed the right agreement. How do you enforce that automatically at every boundary, not just at the front door? AEP answers all three with a single protocol that any runtime can implement, regardless of language or platform. ## How AEP Works The protocol operates on one structural principle: **context flows down, results flow up.** When an agent calls a sub-agent, it passes down an execution context carrying identity, budget, governance rules, and tracing information. When the sub-agent finishes, it returns an execution envelope containing results, cost records, citations, and audit trails. This composes recursively. In a three-organization workflow, each boundary adds a layer of context on the way down and a layer of accountability on the way up. By the time results reach the root caller, there's a complete tree of who spent what, who produced which conclusion, and who was authorized to see which data. ## Three Capabilities **Cost accountability.** Costs form a tree mirroring the execution tree. Every participant records its own costs and includes its children's costs. Budget limits flow down and are enforced at every level. The root caller sees the complete breakdown without needing access to any participant's internal systems. **Provenance.** Every conclusion can carry citations back to the data, model, and prompt that produced it. Citations compose across organizational boundaries -- footnotes that go all the way to primary sources, even when three different organizations produced the intermediate analyses. **Data governance.** Data carries classification labels (public, confidential, PHI, PII) and consent is checked at every organizational boundary. An intermediary can't pass sensitive data to a downstream processor that hasn't been authorized. Every governance decision is recorded in an audit trail designed for GDPR, HIPAA, and SOC 2 compliance. ## Incremental Adoption AEP defines four conformance levels so organizations can adopt incrementally: | Level | Name | What It Adds | | ----- | ----------- | ---------------------------------------------------------- | | 0 | Minimal | Accept and forward context; return valid envelopes | | 1 | Traceable | Distributed tracing with timing and parent-child spans | | 2 | Accountable | Cost recording, aggregation, and budget enforcement | | 3 | Governed | Data classification, consent enforcement, and audit trails | A marketplace might require Level 2 for all participants (so costs can be settled) while only requiring Level 3 for workflows involving regulated data. A simple tool wrapper can reach Level 0 with minimal effort. ## Read the Full Specification This page is the starting point. The full AEP specification is a comprehensive document that covers everything from the high-level architecture to the complete proto3 wire format definitions. **[Read the AEP Whitepaper](/docs/aep-whitepaper)** -- the complete protocol specification, structured so you can read as deep as you need: - **Sections 1--5** cover motivation, architecture, and the three capabilities in detail. No prior protocol experience needed. If you're an executive, investor, or product leader, these sections give you the complete picture. - **Sections 6--10** are the full technical specification using RFC 2119 conventions (MUST, SHOULD, MAY). Proto3 message definitions, wire format rules, executor contracts, and conformance requirements. This is what you need to build an AEP-compatible runtime. - **Appendices** consolidate the complete proto3 schema, collected JSON examples from the running three-organization scenario, and a glossary. A running example -- three organizations collaborating on a clinical data analysis under HIPAA constraints -- threads through every section, so abstract concepts are always grounded in a concrete scenario. ## Design Principles - **Vendor-neutral.** A protocol specification, not a platform feature. Any runtime that implements it can participate. - **Language-agnostic.** Defined using Protocol Buffers with JSON encoding. Implementations can be built in any language. - **Composable.** The same context/envelope pattern governs agents, workflows, tools, and remote calls. - **Incrementally adoptable.** Conformance levels let participants start simple. Workflow extensions are purely additive. - **Open.** Apache 2.0 licensed. A public good for the ecosystem. --- ## Agent Compute Protocol Source: https://aceteam.ai/docs/acp-overview The protocol that lets AI agents dynamically request compute resources at runtime. ACP is the compute substrate counterpart to AEP's accountability layer. # ACP: Agent Compute Protocol AI agents need compute. A document indexing workflow needs 5 workers to process 100K documents. A multi-model analysis needs GPU capacity for 10 minutes. A parallel data pipeline needs to burst to 20 concurrent workers, then scale back to zero. Traditionally, compute is statically provisioned: you decide how many workers to run, and agents work within those limits. ACP inverts this. Agents request capacity at runtime, the platform checks availability and budget atomically, and returns a queue endpoint for job submission. ACP is the counterpart to AEP. Where AEP handles accountability -- costs, citations, governance -- ACP handles the **compute substrate**: how agents acquire, use, and release the workers that do the actual work. ## Why This Matters Static provisioning wastes resources or creates bottlenecks. You either over-provision (paying for idle workers) or under-provision (agents queue behind each other). Neither scales. ACP solves this with a simple negotiation: 1. **Request.** An agent says: "I need 5 CPU workers for 10 minutes, budget $15." 2. **Check.** The platform checks capacity and budget atomically -- no race conditions. 3. **Allocate.** Workers are reserved and a queue endpoint is returned. 4. **Use.** The agent submits jobs to the queue. Workers process them. 5. **Release.** The agent releases workers when done, and receives a usage report. ## The Four Messages ACP defines four messages between agents and the Resource Broker: | Message | Direction | Purpose | | -------------- | --------------- | --------------------------------------------------------------- | | `allocate` | Agent -> Broker | Request workers by capability, count, duration, and budget | | `allocated` | Broker -> Agent | Confirm reservation with allocation ID, queue endpoint, pricing | | `release` | Agent -> Broker | Return workers and end the allocation | | `usage_report` | Broker -> Agent | Final cost, duration, and worker-seconds consumed | ### Allocation Statuses | Status | Meaning | | --------- | --------------------------------------------------------- | | `granted` | Full request fulfilled -- all requested workers allocated | | `partial` | Some workers allocated, fewer than requested | | `denied` | No workers available or budget insufficient | | `queued` | Request accepted, waiting for capacity (planned) | ## Three-Level Scaling ACP is designed as a progressive scaling system: | Level | Mechanism | What It Does | Status | | ----- | ------------------- | -------------------------------------------------------------------------------------- | ----------- | | **1** | Static pool + Redis | Workers self-register capacity on connect. Broker reserves atomically via Lua scripts. | Implemented | | **2** | KEDA autoscaling | When capacity hits zero, KEDA scales up new worker pods based on queue depth. | Planned | | **3** | KAI Scheduler | GPU-aware scheduling for heterogeneous hardware (A100, H100, RTX 3090). | Future | At Level 1, workers announce their capabilities when they connect (`HINCRBY worker:capacity:{type} +1`). The Resource Broker uses atomic Redis Lua scripts to check and reserve capacity in a single operation -- no race conditions between concurrent allocation requests. ## How ACP and AEP Work Together AEP and ACP operate at different layers but share the same execution context: - **AEP** tracks _what happened_: costs, citations, governance decisions. - **ACP** manages _where it happens_: compute allocation, scaling, resource lifecycle. When an agent uses ACP to allocate workers and AEP to track costs, the platform has complete visibility: which resources were used, for how long, at what cost, and who authorized the spend. ## Design Principles - **Agent-initiated.** Agents request resources; the platform fulfills or denies. No static provisioning required. - **Budget-aware.** Every allocation request includes a budget. The broker enforces spending limits before reserving capacity. - **Atomic.** Capacity checks and reservations happen in a single Redis Lua script. No TOCTOU race conditions. - **Composable.** Works with any job queue system. Level 1 uses Redis Streams; higher levels can use Kubernetes-native scheduling. - **Open.** Apache 2.0 licensed alongside AEP. --- ## AEP Whitepaper Source: https://aceteam.ai/docs/aep-whitepaper The complete Agentic Execution Protocol specification: cost accountability, provenance, and data governance for multi-organization AI workflows. # Agentic Execution Protocol (AEP) **Status:** Draft v0.1 **Date:** 2026-02-04 **Authors:** AceTeam Engineering **License:** Apache 2.0 --- ## Abstract Organizations are connecting AI agents into workflows that cross organizational boundaries, but three critical gaps prevent these workflows from operating with the accountability that professional services have always required: there are no receipts (cost accountability), no footnotes (provenance), and no NDAs (data governance). The Agentic Execution Protocol (AEP) closes these gaps with a vendor-neutral wire protocol that defines how execution context flows downward through a call tree and how results, costs, citations, and governance metadata flow upward. AEP is language-agnostic (defined in Protocol Buffers with JSON encoding), composable (the same contract governs agents, workflows, tools, and remote calls), and incrementally adoptable (four conformance levels from basic tracing to full governance). Any runtime that speaks AEP can participate in cross-organization agent workflows with full accountability. ## Table of Contents **Part I: Motivation and Architecture** (Non-Technical) - [1. Introduction](#1-introduction) - [2. Protocol Architecture](#2-protocol-architecture) - [3. Cost Accountability](#3-cost-accountability) - [4. Provenance and Observability](#4-provenance-and-observability) - [5. Data Governance](#5-data-governance) **Part II: Protocol Specification** (Technical: RFC-Style) - [6. Core Messages](#6-core-messages) - [7. Executor Contract](#7-executor-contract) - [8. Wire Protocols](#8-wire-protocols) - [9. Conformance](#9-conformance) **Part III: Context and Future** - [10. Roadmap](#10-roadmap) **Appendices** - [Appendix A: Complete Proto3 Schema](#appendix-a-complete-proto3-schema) - [Appendix B: Collected JSON Examples](#appendix-b-collected-json-examples) - [Appendix C: Relationship to Existing Systems](#appendix-c-relationship-to-existing-systems) - [Appendix D: Glossary](#appendix-d-glossary) --- **Conventions note:** Sections 1–5 use plain language accessible to non-technical readers. Starting in Section 6, this document uses RFC 2119 keywords ("MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL") as defined in [RFC 2119](https://www.rfc-editor.org/rfc/rfc2119), to specify protocol requirements precisely. # Part I: Motivation and Architecture ## 1. Introduction ### 1.1 The Problem Professional services have always been recursive. A management consulting firm subcontracts a market research specialist, who subcontracts a data analytics firm. Money flows down the chain. Findings flow back up. At every step, there are receipts documenting what was spent, footnotes attributing conclusions to their sources, and NDAs ensuring confidential data doesn't reach unauthorized parties. AI agent workflows have the same structure but none of the paperwork. A consulting firm's strategy agent calls a research firm's analysis agent, which calls an analytics firm's data processing agent. Three organizations, three runtimes, one workflow. Yet each platform tracks costs internally with no receipt that travels with the request. There is no citation chain linking a conclusion back to the sub-agent that produced it. There is no consent mechanism preventing sensitive data from reaching an agent whose organization hasn't signed the right agreement. The gaps are structural: - **No receipts.** Cost tracking is flat. When Org C pays for a workflow that touches three organizations, Org C sees one number, not a breakdown by organization, by step, or by resource. - **No footnotes.** When an agent produces a conclusion, there is no way to trace which sub-agent produced which finding, which model was used, or what data was consulted. - **No NDAs.** Data classification exists in policy documents but not in the wire protocol. Sensitive data can flow to downstream agents whose organizations lack the appropriate consent. ### 1.2 Running Example Throughout this document, we use a single scenario to make every concept concrete: > **Org C** ("Apex Consulting") engages **Org A**'s ("ResearchBot Inc") research agent via a marketplace. Org A's agent internally invokes **Org B**'s ("MedExtract Co") document extraction workflow to process clinical data. The execution crosses two organizational boundaries, involves Protected Health Information (PHI) subject to HIPAA, and must track costs for settlement between all three parties. This is not hypothetical. It represents a common pattern in healthcare consulting where AI-assisted clinical document review involves multiple specialized providers. ### 1.3 What AEP Is AEP is the shared paperwork format for AI agent workflows. It is a vendor-neutral wire protocol that defines three capabilities: - **Cost accountability**: the receipts. Hierarchical cost trees that mirror the execution tree, with budget enforcement at every level. - **Provenance**: the footnotes. Citations that trace any claim back to its source agent, model, and data, composing across organizational boundaries. - **Data governance**: the NDAs. Field-level data classification with consent-based boundary enforcement and a complete audit trail. AEP is not a framework, a library, or a platform feature. It is a protocol specification defined using Protocol Buffers with JSON as the default encoding. Any runtime (in Go, Python, Rust, TypeScript, or any language with protobuf support) can implement AEP and participate in cross-organization workflows with full accountability. ## 2. Protocol Architecture ### 2.1 Bidirectional Flow AEP operates on a structural principle: context flows down, results flow up. When an agent calls a sub-agent, whether within the same organization or across a boundary, an **execution context** is constructed and passed to the callee. This context carries the caller's identity, a budget allocation, tracing information, prompt configuration, and governance policies that constrain what the callee can do with the data it receives. When the callee finishes, it returns an **execution envelope** containing its results. The envelope carries more than the output: a cost record documenting what was spent (and what sub-callees spent), provenance citations linking conclusions to their sources, governance audit records documenting data access decisions, and tracing spans for the execution tree. This bidirectional flow composes recursively: ``` context flows down ┌───────────────┐ │ ▼ ┌──────────┐ │ ┌──────────────┐ ┌──────────────┐ │ Org C │─────┘ │ Org A │────▶│ Org B │ │ (caller) │◀──────────│ (specialist) │◀────│(sub-specialist)│ └──────────┘ └──────────────┘ └──────────────┘ │ ▲ └───────────────┘ results + costs flow up ``` In our running example: Org C (Apex Consulting) calls Org A (ResearchBot Inc) with a $50 budget and governance policy requiring HIPAA compliance. Org A calls Org B (MedExtract Co) with a sub-budget and the same governance policy. Org B executes its extraction workflow, records costs and citations, and returns an envelope. Org A wraps that envelope with its own costs and citations and returns the combined result to Org C. Org C now has a complete picture: total cost broken down by organization, a citation chain from the final clinical summary to the source extraction, and an audit trail of every data governance decision made at each boundary. ### 2.2 The Executor Contract The fundamental abstraction in AEP is the **executor**: any component that accepts an execution context and returns an execution envelope. ``` Execute(ExecutionContext) → (ExecutionEnvelope, error) ``` This uniform interface governs all execution types: - **Agent**: An LLM-powered reasoning loop that iterates through tool calls. - **Workflow**: A DAG-based orchestrator that schedules nodes in parallel. - **Tool**: A single operation (HTTP call, database query, model inference). - **Remote**: A cross-organizational call that serializes context over a network boundary. Composability follows naturally. A workflow node can be an agent. An agent's tool can be a workflow. A tool can be a remote call to another organization's service. All use the same contract. In our running example, the full call tree is: ``` AgentExecutor (Org A, research agent) ├── ToolExecutor (Org A, LLM call #1) ├── RemoteExecutor → WorkflowExecutor (Org B, extraction workflow) │ ├── ToolExecutor (Org B, text chunking) │ ├── ToolExecutor (Org B, GLiNER2 entity extraction) │ └── ToolExecutor (Org B, GPT-4o structured output) └── ToolExecutor (Org A, LLM call #2 — final response) ``` Every node in this tree implements the same `Execute(context) → envelope` contract. ### 2.3 Conformance Levels AEP defines four conformance levels to support incremental adoption: | Level | Name | What It Requires | | ----- | ---------- | ------------------------------------------------------------------------ | | 0 | Minimal | Accept and return execution contexts and envelopes with basic tracing | | 1 | Cost-Aware | Level 0 + record and aggregate costs; enforce budget limits | | 2 | Governed | Level 1 + enforce data classification and consent; produce audit records | | 3 | Full | Level 2 + all executor types, prompt layers, all wire protocols | Each level builds on the previous one. A marketplace might require Level 1 for all participants (so costs can be settled) while only requiring Level 2 for workflows involving regulated data like PHI. This tiered approach serves two purposes. First, it lowers the barrier to entry: a simple tool wrapper can reach Level 0 with minimal effort. Second, it provides a clear vocabulary for trust. When an organization advertises Level 2 conformance, that is a testable protocol guarantee, not a marketing claim. ### 2.4 Design Principles **Vendor-neutral.** AEP is a protocol specification, not a library or platform feature. Any runtime that implements the protocol can participate in AEP workflows, regardless of language, platform, or deployment model. **Language-agnostic.** The protocol is defined using Protocol Buffers with JSON as the default encoding. Implementations can be built in any language with protobuf support. **Composable.** The same protocol governs agents, workflows, tools, and remote service calls. An agent calling a tool uses the same context/envelope pattern as a workflow calling a sub-workflow across organizational boundaries. **Incrementally adoptable.** Conformance levels allow participants to start simple and add capabilities over time. Workflow extensions are purely additive: existing workflows don't break. **Open.** AEP is Apache 2.0 licensed. The protocol specification is a public good intended to benefit the ecosystem, not lock participants into any single platform. ## 3. Cost Accountability ### 3.1 Hierarchical Cost Trees Costs in AEP form a tree that mirrors the execution tree. Every executor (whether an agent, a workflow node, or a tool invocation) records what it spent and what its children spent. When an agent calls three sub-agents, its cost record contains its own direct costs plus the cost records from each sub-agent, which in turn contain their own sub-costs. Each cost node tracks three components: | Component | Description | | --------------- | ------------------------------------------------------- | | Compute cost | Raw infrastructure cost: LLM API charges, GPU time, I/O | | Value-added fee | The executor organization's markup for their service | | Platform fee | Platform commission on marketplace transactions | In our running example, the cost tree breaks down as: ``` Org A's research agent (called by Org C) ├── Compute cost: $5.00 (LLM reasoning) ├── Value-added fee: $5.00 (Org A's markup) ├── Platform fee: $0.75 (5% of $10.00 + $5.25) └── Org B's extraction workflow (called by Org A) ├── Compute cost: $3.00 (GLiNER2 $1.50 + GPT-4o $1.50) ├── Value-added fee: $2.00 (Org B's markup) └── Platform fee: $0.25 (5% of $5.00) Total bill to Org C: $16.00 ``` The root caller (Org C) sees the complete breakdown. No single platform needs global visibility. Each participant records its own costs honestly, and the protocol aggregates them as results flow back up. ### 3.2 Budget Enforcement Budget limits flow downward with the context. A caller can say "this sub-call has a budget of $50.00." The callee is responsible for not exceeding that budget and for further subdividing it among its own sub-calls. AEP uses **pessimistic reservation** to prevent overspending: 1. **Reserve.** Before making any call that incurs cost, the executor estimates the worst-case cost (e.g., `max_tokens * price_per_token` for an LLM call) and reserves that amount from the available budget. If the budget is insufficient, execution stops immediately. 2. **Execute.** The operation proceeds. 3. **Settle.** After the operation completes, the reservation is released and the actual cost is recorded. The difference (reservation minus actual) returns to the available budget. This model is analogous to credit card authorization-then-capture. It ensures that budget is never exceeded, even when multiple operations execute concurrently. When execution crosses an organizational boundary, the caller allocates a sub-budget to the callee. The callee never sees the caller's total budget, only its allocated portion. This prevents information leakage while maintaining budget discipline. ### 3.3 Settlement AEP separates cost recording from settlement. The cost tree captures what each organization spent and earned. How those amounts are settled is determined by a pluggable adapter: - **Phase 1:** An internal ledger adapter implements double-entry bookkeeping in a database: instant settlement, zero fees. - **Future phases:** Blockchain adapters (Solana, Ethereum, Cardano) bridge the internal ledger to on-chain settlement for organizations that require it. The platform takes a commission (e.g., 5%) on marketplace transactions, recorded as a platform fee in each cost node. Settlement timing is configurable: jobs complete at different times, so batching at regular intervals is natural. The key insight is that cost accountability is structural, not administrative. It doesn't require a central billing system. It requires only that each participant follows the protocol: record your costs, respect your budget, and include your children's costs in your envelope. ## 4. Provenance and Observability ### 4.1 Citations Every output an AEP executor produces can carry citations back to the inputs, models, and reasoning that produced it. Citations are the AI equivalent of footnotes in a research paper. Agents embed citation references in their output using a marker format: `[ref:span-id]`. The framework resolves these markers to full citation objects in the envelope, allowing downstream consumers to: 1. Display inline citations in the UI. 2. Trace any claim back to its source span, organization, and data classification. 3. Apply governance rules (redaction, consent checks) to specific parts of the output. In our running example, when Org A's research agent produces a clinical summary, it cites the extraction results from Org B: > "The patient presents with Type 2 Diabetes Mellitus (ICD-10: E11.9) with peripheral neuropathy (G63). HbA1c: 8.2% (above target) **[ref:span-org-b-extraction-1234]**" The citation chain goes from Org A's conclusion to Org B's extraction workflow, which itself can cite the specific GLiNER2 entity extraction model that identified the diagnosis. This gives the end user (Org C) a complete provenance trail. Citations compose across organizational boundaries just like costs. When Org A cites a finding from Org B, and Org B's workflow cites specific data sources, the final chain traces from Org A's conclusion all the way to the original clinical document. ### 4.2 Distributed Tracing Every executor creates a **span**: a structured record of what happened, how long it took, and how it relates to parent and child executions. Spans form a tree rooted at a single trace ID. In our running example, the span tree is: ``` trace-a1b2c3d4e5f6 └── span-org-c-root-0001 (Org C, agent, root) └── span-org-a-agent-7890 (Org A, agent, marketplace call) ├── span-org-a-llm-01 (Org A, tool, LLM call #1) ├── span-org-b-extraction-1234 (Org B, workflow, marketplace call) │ ├── span-org-b-chunker-01 (Org B, tool, text chunking) │ ├── span-org-b-gliner-5678 (Org B, tool, GLiNER2 extraction) │ └── span-org-b-llm-9012 (Org B, tool, GPT-4o structured output) └── span-org-a-llm-02 (Org A, tool, LLM call #2 — final response) ``` Each span carries timing data, an executor type, an organization ID, and status. Spans carry organizational visibility rules: when returning results across an org boundary, the callee controls which internal spans are exposed to the caller. ### 4.3 Regulatory Compliance Together, citations and tracing answer two questions that regulators, auditors, and customers increasingly ask: **"How did the AI reach this conclusion?"** The citation chain traces any claim in the final output back to its source data, model, and organization. For clinical data under HIPAA, this means every finding can be linked to the specific extraction model, the specific document, and the specific patient record. **"What happened during this execution?"** The span tree provides a complete execution trace: which organizations were involved, what operations they performed, how long each took, and whether any errors occurred. This satisfies audit requirements for GDPR (Article 30 processing records), HIPAA (audit controls), and SOC 2 (monitoring). ## 5. Data Governance ### 5.1 Field-Level Classification Data governance in AEP operates at the field level, not the document level. A patient record might have PHI fields (name, date of birth) and non-PHI fields (visit count, facility name). AEP classifies each field independently using dot-notation paths: | Security Level | Description | | -------------- | ------------------------------------------------ | | PUBLIC | No restrictions | | INTERNAL | Visible within the owning organization only | | CONFIDENTIAL | Restricted access, requires authorization | | PHI | Protected Health Information (HIPAA) | | PII | Personally Identifiable Information (GDPR, CCPA) | Each classification carries constraints that travel with the data: - **Redistribution**: can the receiving organization pass this data to a third party? - **Encryption**: must the data be encrypted at rest? - **Retention**: how long may the data be kept? (e.g., "7y" for HIPAA, "30d" for analytics) - **Purposes**: for what purposes may the data be used? (e.g., "treatment", "billing") - **Redaction**: must the field be redacted before crossing an organizational boundary? - **Audit**: must all access be logged for compliance? In our running example, Org B's extraction workflow produces entities from clinical documents. The extracted HbA1c value (8.2%) is classified as PHI with constraints: no redistribution, encryption required, 7-year retention, allowed purposes limited to treatment and billing, audit required. ### 5.2 Consent-Based Boundary Enforcement Before data crosses an organizational boundary, AEP checks whether consent exists for that transfer. A **consent grant** records a specific authorization, for example, "Patient uuid-pat-5678 has authorized Org A (ResearchBot Inc) to access PHI for treatment and billing purposes, valid until 2027-02-04." Enforcement is recursive. If Org A calls Org B, and Org B calls Org C, the governance check happens at both boundaries independently. Org C doesn't inherit Org A's consent grants. It must have its own. This prevents a common vulnerability where a trusted intermediary inadvertently passes sensitive data to an unauthorized downstream processor. When a consent check fails, the system handles it according to the classification constraints: - If the field is marked as requiring redaction: the field is redacted from the result before crossing the boundary. - Otherwise: the entire response is rejected with a consent-denied error. Every consent check, both allowed and denied, is logged as an audit record. ### 5.3 Audit Trail Every governance decision flows back up with the results in the execution envelope. The root caller (Org C) can see every consent check, every data classification decision, and every boundary crossing that occurred during the entire execution tree. In our running example, Org C's audit trail shows: - Org B classified extracted entities as PHI with HIPAA regulations. - Consent was verified at the Org B → Org A boundary for patient uuid-pat-5678, purposes: treatment and billing. - Consent was verified at the Org A → Org C boundary for the same patient. - All access was logged with audit record IDs for compliance retrieval. This audit trail is designed to satisfy regulatory requirements: GDPR Article 30 (records of processing activities), HIPAA audit controls, and SOC 2 monitoring requirements. ### 5.4 Consent Revocation When a consent grant is revoked, AEP defines a cascade procedure: 1. **Query affected spans.** Identify all spans where the data subject was accessed, using the data classification's subject identifiers. 2. **Identify downstream outputs.** Find cached results, stored extractions, and derived outputs that incorporate the affected data. 3. **Mark stale.** Flag affected cached results as stale or trigger re-evaluation without the revoked data. 4. **Audit.** Create an audit record documenting the cascade: what was affected, what actions were taken. # Part II: Protocol Specification > Sections 6–10 use RFC 2119 keywords precisely. All message definitions use proto3 syntax and are language-agnostic. Examples use JSON encoding. ## 6. Core Messages ### 6.1 Encoding All messages in AEP are defined using Protocol Buffers v3 (proto3). Two wire encodings are supported: - **JSON encoding**: for REST APIs, Redis Streams, and human-readable debugging. Field names use `snake_case` per proto3 JSON mapping rules. Implementations MUST support JSON encoding. - **Binary protobuf encoding**: for gRPC, fabric mesh inter-node calls, and performance-sensitive paths. Binary protobuf encoding is OPTIONAL. ### 6.2 ExecutionContext The `ExecutionContext` carries identity, budget, tracing, prompt configuration, and governance policy downward through the execution tree. Every executor receives an `ExecutionContext` and derives a child context before calling sub-executors. ```protobuf syntax = "proto3"; package aep.core.v1; import "google/protobuf/any.proto"; import "google/protobuf/struct.proto"; import "google/protobuf/timestamp.proto"; message ExecutionContext { repeated OrgIdentity caller_chain = 1; // ordered: [originator, ..., immediate caller] string project_id = 2; // resource scoping within an organization string user_id = 3; // originating end-user BudgetState budget = 4; // remaining budget and reservations string trace_id = 5; // root trace ID (shared across entire execution tree) string parent_span_id = 6; // parent span in execution tree string span_id = 7; // this node's span repeated PromptLayer prompt_stack = 8; // ordered: global -> org -> agent GovernanceContext governance_context = 9; map values = 10; // extensibility } message OrgIdentity { string org_id = 1; string org_name = 2; // human-readable, for logging only repeated string roles = 3; // caller's roles in this org } ``` **Running example:** When Org B's extraction workflow receives its context, `caller_chain` contains three entries: Org C (originator), Org A (intermediate caller), Org B (current executor). The `trace_id` is shared across all three organizations, linking the entire execution tree. **Context derivation rules.** When an executor calls a sub-executor, it MUST derive a child context: 1. Append its own `OrgIdentity` to `caller_chain` if an org boundary is being crossed. 2. Set `parent_span_id` to the current `span_id`. 3. Generate a new unique `span_id`. 4. Copy `trace_id` unchanged. 5. Copy `budget` by reference for in-process calls, or allocate a sub-budget for cross-boundary calls. 6. Append any org-specific or agent-specific `PromptLayer` entries. 7. Copy `governance_context` unchanged. ### 6.3 ExecutionEnvelope The `ExecutionEnvelope` carries results, cost trees, execution spans, citations, and errors upward from an executor to its caller. ```protobuf message ExecutionEnvelope { google.protobuf.Any result = 1; // the executor's output (type varies by executor) CostNode cost_tree = 2; // hierarchical cost breakdown repeated Span spans = 3; // execution trace entries repeated Citation citations = 4; // provenance and source attribution repeated ExecError errors = 5; // non-fatal errors and warnings } message ExecError { string code = 1; // machine-readable error code string message = 2; // human-readable description string span_id = 3; // span where error occurred Severity severity = 4; enum Severity { WARNING = 0; ERROR = 1; FATAL = 2; } } ``` **Envelope merging rules.** When an executor receives envelopes from multiple sub-executors (e.g., a workflow executing parallel nodes), it MUST merge them: 1. Collect all `spans` from child envelopes into the parent envelope's `spans` array. 2. Attach child `cost_tree` nodes as `children` of the parent's `CostNode`. 3. Collect all `citations` from child envelopes. 4. Collect all `errors` from child envelopes. 5. The `result` is determined by the parent executor's logic (e.g., the agent's final LLM response, or the workflow's terminal node output). ### 6.4 CostNode The `CostNode` is a recursive tree structure that mirrors the execution tree. ```protobuf message CostNode { string span_id = 1; // links to the corresponding Span string executor_org = 2; // organization that performed the work string caller_org = 3; // organization that requested the work string service_type = 4; // "agent", "workflow", "tool", "remote" string service_id = 5; // marketplace service identifier string compute_cost = 6; // raw infrastructure cost (decimal string) string value_added_fee = 7; // executor org's markup (decimal string) string platform_fee = 8; // platform commission (decimal string) repeated CostNode children = 9; // sub-executor costs } ``` **TotalCost formula.** The total cost of a `CostNode` is defined recursively: ``` TotalCost(node) = node.compute_cost + node.value_added_fee + node.platform_fee + sum(TotalCost(child) for child in node.children) ``` **Fee structure:** | Fee Type | Description | Determined By | | ----------------- | --------------------------------------------------------------- | -------------------------------------------------------------- | | `compute_cost` | Raw infrastructure cost: LLM API charges, GPU time, storage I/O | Executor, based on actual resource consumption | | `value_added_fee` | The executor organization's markup for their service | Executor's pricing model (fixed, per-token, or percentage) | | `platform_fee` | Platform commission on marketplace transactions | Platform policy (e.g., 5% of `compute_cost + value_added_fee`) | All monetary values are represented as decimal strings. Implementations MUST use arbitrary-precision decimal arithmetic for cost calculations to avoid floating-point precision errors. **Running example:** Intra-org calls (Org B calling its own tools) have zero `value_added_fee` and zero `platform_fee`. Cross-org calls (Org A calling Org B) incur all three fee types. ### 6.5 BudgetState `BudgetState` tracks the financial budget for an execution tree using pessimistic reservation. ```protobuf message BudgetState { string total = 1; // total budget authorized by root caller (decimal string) string spent = 2; // accumulated actual cost (decimal string) string reserved = 3; // pessimistic reservations for in-flight calls (decimal string) } ``` **Available budget formula:** ``` Available(budget) = budget.total - budget.spent - budget.reserved ``` **Reserve/Settle pseudocode:** ``` Reserve(amount): if Available() < amount: return BUDGET_EXCEEDED reserved += amount return Reservation{amount} Settle(reservation, actual_cost): reserved -= reservation.amount spent += actual_cost ``` **Cross-boundary allocation rules.** When execution crosses an org boundary: 1. The caller MUST NOT expose its total budget to the callee. 2. The caller allocates a sub-budget: `sub_total = min(estimated_cost * safety_factor, Available())`. 3. The callee receives a `BudgetState` with `total = sub_total, spent = 0, reserved = 0`. 4. When the callee returns, its `spent` is settled against the caller's reservation. ### 6.6 Span A `Span` represents a single unit of work in the execution tree. ```protobuf message Span { string span_id = 1; string parent_span_id = 2; // empty for root spans string trace_id = 3; // root trace ID string executor_type = 4; // "agent", "workflow", "tool", "remote" string executor_id = 5; string org_id = 6; // organization that owns this executor google.protobuf.Timestamp start_time = 7; google.protobuf.Timestamp end_time = 8; SpanStatus status = 9; map metadata = 10; } enum SpanStatus { OK = 0; ERROR = 1; CANCELLED = 2; BUDGET_EXCEEDED = 3; } ``` Spans form a tree via `parent_span_id` references, rooted at a single `trace_id`. **Visibility rules.** Spans carry an `org_id`. When returning an `ExecutionEnvelope` across an org boundary: - The calling org SHOULD receive spans for the boundary executor (the top-level span of the callee). - The calling org MAY receive sub-spans depending on the callee's visibility policy. - Implementations SHOULD support a visibility policy on the executor that controls which internal spans are exposed to callers. ### 6.7 Citation A `Citation` attributes part of an executor's output to a specific source. ```protobuf message Citation { string span_id = 1; // the span that produced this cited content string source_org = 2; // organization that owns the source string source_type = 3; // "agent", "workflow", "retrieval", "tool", "document" string source_id = 4; string content = 5; // the relevant excerpt or summary double confidence = 6; // confidence score (0.0-1.0) repeated DataClassification classifications = 7; // data governance tags } ``` **Reference marker format.** Agents SHOULD embed citation references in their output using the marker `[ref:span-id]`. The framework resolves these markers to full `Citation` objects in the envelope. **Running example:** Org A's research agent outputs: "HbA1c: 8.2% (above target) [ref:span-org-b-extraction-1234]". The framework resolves this to a Citation object linking to Org B's extraction span, with PHI classification and a confidence score of 0.94. ### 6.8 DataClassification and Consent **SecurityLevel enum:** ```protobuf enum SecurityLevel { PUBLIC = 0; INTERNAL = 1; CONFIDENTIAL = 2; PHI = 3; PII = 4; } ``` **DataClassification and DataConstraints:** ```protobuf message DataClassification { SecurityLevel level = 1; repeated string regulations = 2; // ["HIPAA", "GDPR", "SOC2", "CCPA"] repeated string data_subjects = 3; // ["patient:uuid-123", "cohort:clinic-xyz"] string field_path = 4; // dot-notation: "patient.name", "entities[0].dob" string source_span_id = 5; // span where classification was assigned DataConstraints constraints = 6; } message DataConstraints { bool allow_redistribution = 1; bool requires_encryption = 2; string retention_policy = 3; // "30d", "7y", "session-only", "indefinite" repeated string allowed_purposes = 4; // ["treatment", "billing", "research"] bool redaction_required = 5; bool audit_required = 6; } ``` **GovernanceContext:** ```protobuf message GovernanceContext { string consent_endpoint = 1; repeated ClassificationPolicy classification_policies = 2; } message ClassificationPolicy { string org_id = 1; SecurityLevel default_level = 2; bool phi_detection_enabled = 3; } ``` **ConsentGrant and ConsentDecision:** ```protobuf message ConsentGrant { string grant_id = 1; string data_subject = 2; // "patient:uuid-123" string granted_to = 3; // "org:org-a-uuid" repeated string purposes = 4; // ["treatment", "billing"] repeated string fields = 5; // specific fields; ["*"] = all google.protobuf.Timestamp expiry = 6; bool revocable = 7; string granted_by = 8; google.protobuf.Timestamp granted_at = 9; } message ConsentDecision { bool allowed = 1; string reason = 2; repeated string constraints = 3; // conditions applied google.protobuf.Timestamp expires_at = 4; string audit_record_id = 5; // every check is logged } ``` **Enforcement algorithm.** When an `ExecutionEnvelope` crosses an organizational boundary, the boundary executor MUST: 1. Walk all `Citations` in the envelope. 2. For each citation with non-empty `classifications`, call the consent service for the receiving organization. 3. If `ConsentDecision.allowed` is `false`: - If `constraints.redaction_required` is `true` on the classification: redact the classified fields from `result`. - Otherwise: reject the entire response with an `ExecError` of severity `FATAL` and code `CONSENT_DENIED`. 4. Log an audit record for every consent check (both allowed and denied). 5. Propagate any `ConsentDecision.constraints` upward in the envelope's `errors` array as warnings. **Revocation cascade.** When a consent grant is revoked: 1. Query all spans where the `data_subject` was accessed (via `DataClassification.data_subjects`). 2. Identify downstream outputs (cached results, stored extractions) incorporating that data. 3. Mark affected cached results as stale or trigger re-evaluation. 4. Create an audit record documenting the cascade. ### 6.9 PromptLayer A `PromptLayer` represents a single layer in a composable prompt stack. ```protobuf message PromptLayer { string source = 1; // "global", "org:", "agent:" string role = 2; // "policy", "persona", "task" string content = 3; // the prompt text int32 priority = 4; // higher value = higher authority } ``` **Source formats:** | Source Pattern | Description | | ------------------------ | --------------------------------------------------------------- | | `global` | Platform-wide policy, applied to all executions | | `org:` | Organization-specific persona or policy | | `agent:` | Agent-specific task instructions | | `workflow:` | Workflow-level instructions for agent nodes within the workflow | **Composition rules:** 1. Layers are ordered by `priority` descending when composed into the LLM's system prompt. 2. A higher-priority layer's instructions MUST NOT be contradicted by a lower-priority layer. If a conflict is detected, the higher-priority layer wins. 3. All layers are visible to the LLM. The implementation SHOULD compose them with clear delimiters so the LLM can distinguish layers. 4. When crossing an org boundary, the callee MAY add its own layers but MUST NOT remove or modify existing layers. **Running example:** Org B's extraction workflow receives a prompt stack with three layers: a global platform policy (priority 100), Org A's research persona (priority 50), and an agent-specific extraction task (priority 10). Org B cannot modify the global policy or Org A's persona, it can only add its own specialization layer. ## 7. Executor Contract ### 7.1 Interface The fundamental interface: ``` Execute(ExecutionContext) → (ExecutionEnvelope, error) ``` In proto3 terms: ```protobuf service ExecutorService { rpc Execute(ExecutionContext) returns (ExecutionEnvelope); rpc ExecuteStream(ExecutionContext) returns (stream StreamEvent); } message StreamEvent { string type = 1; // event type (see Section 8.2) google.protobuf.Timestamp timestamp = 2; google.protobuf.Any payload = 3; // type-specific payload } ``` ### 7.2 Executor Types | Type | Description | Behavior | | ------------ | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | | **agent** | LLM-powered reasoning loop | Iterates: send messages to LLM, parse tool calls, execute tools (each is an Executor), append results, repeat until final response | | **workflow** | DAG-based orchestration | Topologically sorts the DAG, executes ready nodes in parallel, merges envelopes from all nodes | | **tool** | Single operation | HTTP call, database query, model inference, file operation. Produces a single span | | **remote** | Cross-boundary call | Serializes context, sends via gRPC/REST, deserializes envelope, enforces governance at the boundary | ### 7.3 Responsibilities Every executor MUST: 1. Create a `Span` at the start of execution (recording `span_id`, `parent_span_id`, `trace_id`, `start_time`). 2. Complete the `Span` at the end of execution (recording `end_time`, `status`). 3. Include its span(s) in the returned `ExecutionEnvelope.spans`. 4. Reserve budget before incurring cost, and settle after. 5. If status is `BUDGET_EXCEEDED`, return immediately with the envelope containing all spans and costs accumulated so far. Every executor SHOULD: 1. Emit `span_start` and `span_end` stream events if streaming is enabled. 2. Emit `cost` stream events after settling each reservation. 3. Emit `citation` stream events when producing cited content. 4. Emit `budget_warning` stream events when `Available(budget)` drops below a configurable threshold (20% of `total` is RECOMMENDED). ### 7.4 Composition Example The full call tree for our running example, showing how the executor contract composes: ``` AgentExecutor.Execute(ctx) [Org A] |-- creates span-org-a-agent-7890 |-- reserves budget for LLM call |-- calls LLM → gets tool call request |-- settles LLM reservation |-- resolves tool "clinical-extraction" → RemoteExecutor | +-- RemoteExecutor.Execute(ctx.WithNewSpan()) [boundary: Org A → Org B] | |-- creates span-remote | |-- allocates sub-budget ($8.00 from Org A's remaining) | |-- serializes context → gRPC to Org B | | | +-- WorkflowExecutor.Execute(sub-ctx) [Org B] | | |-- creates span-org-b-extraction-1234 | | |-- executes DAG nodes in parallel: | | | +-- ToolExecutor("chunker").Execute(sub-ctx.WithNewSpan()) | | | +-- ToolExecutor("gliner2").Execute(sub-ctx.WithNewSpan()) | | | +-- ToolExecutor("gpt4o").Execute(sub-ctx.WithNewSpan()) | | |-- merges child envelopes | | |-- returns ExecutionEnvelope | | | |-- deserializes envelope | |-- enforces governance (consent checks on citations) | |-- settles sub-budget against caller's reservation | |-- returns merged ExecutionEnvelope | |-- merges tool envelope into agent envelope |-- reserves budget for final LLM call |-- calls LLM with tool results → gets final response |-- settles final LLM reservation |-- returns ExecutionEnvelope with all spans, costs, citations ``` ## 8. Wire Protocols AEP defines three wire protocol contexts: job queues, streaming events, and cross-boundary serialization. ### 8.1 Job Queue AEP extends the existing Redis Streams job queue pattern. The base job payload structure remains unchanged for backward compatibility; AEP fields are added as optional extensions. **Base payload (existing):** ```json { "version": "1.0", "type": "agent_response", "jobId": "job-uuid-1234", "rayId": "ray-SFO-2m5x7k-a1b2c3", "userId": "user-jane-doe-uuid", "organizationId": "org-c-uuid", "createdAt": "2026-02-04T10:00:00.000Z", "priority": 0, "maxAttempts": 3, "timeoutSeconds": 300, "requiredCapabilities": ["llm:gpt-4o"] } ``` **AEP extension field:** AEP-aware jobs include an optional `execution_context` field containing a serialized `ExecutionContext`. Non-AEP workers ignore this field. **Backward compatibility rules:** - Existing job payloads without `execution_context` remain valid. Workers MUST process them as before. - When `execution_context` is absent, AEP-aware workers SHOULD construct a minimal context from base fields: `userId`, `organizationId`, `rayId` (as `trace_id`). - The `rayId` field is superseded by `execution_context.trace_id` when present, but MUST still be populated for backward compatibility with existing monitoring. ### 8.2 Stream Events AEP extends the existing stream event types with five new types: **Existing types (retained):** | Type | Description | | ----------------- | --------------------------------------------------- | | `start` | Job processing has begun | | `chunk` | Incremental output content (e.g., LLM token stream) | | `tool_call_start` | A tool call has begun | | `tool_call_end` | A tool call has completed | | `error` | An error occurred | | `end` | Job processing is complete | **New AEP types:** | Type | Description | Payload | | ---------------- | ----------------------- | ------------------------------------------------- | | `cost` | Incremental cost update | `CostNode` (partial, for the current span) | | `citation` | A citation was produced | `Citation` | | `span_start` | A new span has begun | `Span` (with `end_time` unset) | | `span_end` | A span has completed | `Span` (with `end_time` and `status` set) | | `budget_warning` | Budget is running low | `{ "available": "...", "threshold": "...", ... }` | ### 8.3 Cross-Boundary Serialization Different transports are used depending on context: | Context | Encoding | Transport | | ----------------------------- | --------------- | --------------------------------- | | REST API (Next.js to backend) | JSON | HTTP/1.1 or HTTP/2 | | Redis Streams job queue | JSON | Redis Streams (XADD/XREADGROUP) | | Redis Pub/Sub streaming | JSON | Redis Pub/Sub (PUBLISH/SUBSCRIBE) | | Fabric mesh (inter-node) | Binary protobuf | gRPC over VPN mesh | | Cross-org marketplace calls | Binary protobuf | gRPC over fabric mesh | | SSE to browser | JSON | Server-Sent Events (HTTP) | **Header propagation.** For REST and gRPC transports, tracing headers MUST be propagated: | Header | Description | Example | | -------------------- | -------------------------------- | ----------------------- | | `X-Trace-Id` | Root trace ID | `trace-a1b2c3d4e5f6` | | `X-Span-Id` | Current span ID | `span-org-a-agent-7890` | | `X-Parent-Span-Id` | Parent span ID | `span-org-c-root-0001` | | `X-Caller-Org` | Immediate caller's org ID | `org-c-uuid` | | `X-Budget-Remaining` | Available budget (informational) | `29.65` | For gRPC, these are carried in gRPC metadata with the same key names. ## 9. Conformance ### 9.1 Level 0: Minimal | Requirement | Keyword | | ------------------------------------------------------------------------------- | ------- | | Implement the `Execute(ExecutionContext) → (ExecutionEnvelope, error)` contract | MUST | | Generate a unique `span_id` for each executor invocation | MUST | | Set `parent_span_id` when calling sub-executors | MUST | | Propagate `trace_id` unchanged through the call tree | MUST | | Include all spans in the returned `ExecutionEnvelope.spans` | MUST | | Set `Span.status` to reflect execution outcome | MUST | | Return `result` in the `ExecutionEnvelope` | MUST | | Support JSON encoding of `ExecutionContext` and `ExecutionEnvelope` | MUST | | Merge child envelopes when composing executors | SHOULD | | Emit `span_start` and `span_end` stream events | SHOULD | | Support binary protobuf encoding | MAY | ### 9.2 Level 1: Cost-Aware | Requirement | Keyword | | ---------------------------------------------------------------------- | ------- | | All Level 0 MUST requirements | MUST | | Implement `BudgetState` with `Reserve` and `Settle` operations | MUST | | Use pessimistic reservation before incurring cost | MUST | | Use arbitrary-precision decimal arithmetic for all cost calculations | MUST | | Populate `CostNode` in returned envelopes | MUST | | Abort with `BUDGET_EXCEEDED` when `Available(budget) < reservation` | MUST | | Allocate sub-budgets for cross-boundary calls | MUST | | Emit `cost` stream events after settling reservations | SHOULD | | Emit `budget_warning` stream events when budget is below threshold | SHOULD | | Track `compute_cost`, `value_added_fee`, and `platform_fee` separately | SHOULD | | Support configurable budget warning thresholds | MAY | ### 9.3 Level 2: Governed | Requirement | Keyword | | -------------------------------------------------------------------------------- | ------- | | All Level 1 MUST requirements | MUST | | Support `DataClassification` on citations | MUST | | Support `SecurityLevel` enum values: PUBLIC, INTERNAL, CONFIDENTIAL, PHI, PII | MUST | | Enforce `DataConstraints` at org boundaries | MUST | | Implement consent checking before allowing classified data across org boundaries | MUST | | Log audit records for every consent check (allowed and denied) | MUST | | Redact fields marked `redaction_required` before crossing org boundaries | MUST | | Support `ConsentGrant` and `ConsentDecision` messages | MUST | | Support consent revocation cascades | SHOULD | | Track `data_subjects` across spans for revocation queries | SHOULD | | Emit `citation` stream events with classifications | SHOULD | | Support field-level (dot-notation) classification paths | MAY | | Support custom `regulations` values beyond the predefined set | MAY | ### 9.4 Level 3: Full | Requirement | Keyword | | ------------------------------------------------------------------------- | ------- | | All Level 2 MUST requirements | MUST | | Support all four executor types: agent, workflow, tool, remote | MUST | | Support `PromptLayer` composition with priority-based conflict resolution | MUST | | Support cross-boundary context serialization | MUST | | Support header propagation for tracing | MUST | | Support the AEP job queue extension | MUST | | Support all AEP stream event types | MUST | | Backward-compatible with existing `BaseJobPayload` | MUST | | Support both JSON and binary protobuf encoding | SHOULD | | Support gRPC `ExecutorService` for cross-boundary calls | SHOULD | | Support bidirectional streaming via `ExecuteStream` | MAY | | Support custom executor types beyond the four predefined types | MAY | ### 9.5 Conformance Declaration Implementations declare conformance in their documentation or service metadata: ```json { "protocol": "AEP-Core", "version": "0.1", "conformance_level": 1, "encoding": ["json"], "executor_types": ["agent", "tool"], "extensions": [] } ``` # Part III: Context and Future ## 10. Roadmap ### 10.1 Current Status The AEP protocol specification (AEP-Core and AEP-Workflows) is in draft. The specifications define the complete wire format, executor contracts, conformance levels, and governance model. The first deployment target is the AceTeam sovereign compute fabric: a distributed infrastructure layer where organizations run AI workloads on their own hardware, connected through a WireGuard mesh network. ### 10.2 Near-Term - **Conformance test suite.** A shared set of JSON test fixtures defining `ExecutionContext` inputs and expected `ExecutionEnvelope` outputs for canonical scenarios (single agent call, nested workflow, cross-org invocation, budget exceeded). Implementations process the same inputs and are validated for structural equivalence. - **Reference implementation.** A Go execution runtime implementing AEP is in active development. Go was chosen for its concurrency model (goroutines map naturally to DAG execution), single-binary deployment, and alignment with the existing sovereign compute CLI. - **SDK libraries.** Python and TypeScript libraries for producing and consuming AEP contexts and envelopes, enabling existing agent frameworks to participate in AEP workflows. - **Framework integrations.** Adapters for existing workflow engines and agent frameworks to adopt AEP incrementally. ### 10.3 AEP-Confidence A future specification (AEP-Confidence) will address uncertainty propagation through execution trees: - **Confidence scores** from extraction models (e.g., GLiNER2 softmax outputs, LLM-predicted confidence). - **Model provenance**: which model produced each entity or conclusion, enabling calibration comparison. - **Uncertainty propagation**: when an agent cites entities extracted at 0.94 confidence, that confidence propagates through the citation chain. - **Multi-model fusion**: running the same input through multiple extraction models and merging results with confidence weighting. ### 10.4 Settlement Evolution The settlement adapter is designed for evolution: - **Phase 1:** Internal ledger adapter: Postgres double-entry bookkeeping. Instant settlement, zero fees. - **Phase 5+:** Blockchain adapters (Solana, EVM chains) for organizations that require on-chain settlement. - **Future:** Cross-chain bridging, smart contract escrow for marketplace transactions, optional staking for service quality guarantees. The ACE token economic model (internal unit of account, purchased via Stripe, with organizational balances and per-agent sub-allocations) provides the accounting foundation. The chain adapter abstraction ensures that business logic never touches chain-specific APIs: all settlement flows through a uniform interface regardless of whether the backend is a database or a blockchain. # Appendices ## Appendix A: Complete Proto3 Schema The complete proto3 schema consolidating all messages defined in this specification: ```protobuf syntax = "proto3"; package aep.core.v1; import "google/protobuf/any.proto"; import "google/protobuf/struct.proto"; import "google/protobuf/timestamp.proto"; // --- Enums --- enum SecurityLevel { PUBLIC = 0; INTERNAL = 1; CONFIDENTIAL = 2; PHI = 3; PII = 4; } enum SpanStatus { OK = 0; ERROR = 1; CANCELLED = 2; BUDGET_EXCEEDED = 3; } // --- Core Messages --- message OrgIdentity { string org_id = 1; string org_name = 2; repeated string roles = 3; } message ExecutionContext { repeated OrgIdentity caller_chain = 1; string project_id = 2; string user_id = 3; BudgetState budget = 4; string trace_id = 5; string parent_span_id = 6; string span_id = 7; repeated PromptLayer prompt_stack = 8; GovernanceContext governance_context = 9; map values = 10; } message ExecutionEnvelope { google.protobuf.Any result = 1; CostNode cost_tree = 2; repeated Span spans = 3; repeated Citation citations = 4; repeated ExecError errors = 5; } message ExecError { string code = 1; string message = 2; string span_id = 3; Severity severity = 4; enum Severity { WARNING = 0; ERROR = 1; FATAL = 2; } } // --- Cost --- message BudgetState { string total = 1; string spent = 2; string reserved = 3; } message CostNode { string span_id = 1; string executor_org = 2; string caller_org = 3; string service_type = 4; string service_id = 5; string compute_cost = 6; string value_added_fee = 7; string platform_fee = 8; repeated CostNode children = 9; } // --- Tracing --- message Span { string span_id = 1; string parent_span_id = 2; string trace_id = 3; string executor_type = 4; string executor_id = 5; string org_id = 6; google.protobuf.Timestamp start_time = 7; google.protobuf.Timestamp end_time = 8; SpanStatus status = 9; map metadata = 10; } // --- Provenance --- message Citation { string span_id = 1; string source_org = 2; string source_type = 3; string source_id = 4; string content = 5; double confidence = 6; repeated DataClassification classifications = 7; } // --- Governance --- message DataClassification { SecurityLevel level = 1; repeated string regulations = 2; repeated string data_subjects = 3; string field_path = 4; string source_span_id = 5; DataConstraints constraints = 6; } message DataConstraints { bool allow_redistribution = 1; bool requires_encryption = 2; string retention_policy = 3; repeated string allowed_purposes = 4; bool redaction_required = 5; bool audit_required = 6; } message GovernanceContext { string consent_endpoint = 1; repeated ClassificationPolicy classification_policies = 2; } message ClassificationPolicy { string org_id = 1; SecurityLevel default_level = 2; bool phi_detection_enabled = 3; } message ConsentGrant { string grant_id = 1; string data_subject = 2; string granted_to = 3; repeated string purposes = 4; repeated string fields = 5; google.protobuf.Timestamp expiry = 6; bool revocable = 7; string granted_by = 8; google.protobuf.Timestamp granted_at = 9; } message ConsentDecision { bool allowed = 1; string reason = 2; repeated string constraints = 3; google.protobuf.Timestamp expires_at = 4; string audit_record_id = 5; } // --- Prompt --- message PromptLayer { string source = 1; string role = 2; string content = 3; int32 priority = 4; } // --- Service --- service ExecutorService { rpc Execute(ExecutionContext) returns (ExecutionEnvelope); rpc ExecuteStream(ExecutionContext) returns (stream StreamEvent); } message StreamEvent { string type = 1; google.protobuf.Timestamp timestamp = 2; google.protobuf.Any payload = 3; } ``` ## Appendix B: Collected JSON Examples All JSON examples use the running example: Org C (Apex Consulting) → Org A (ResearchBot Inc) → Org B (MedExtract Co) processing clinical data. ### B.1 ExecutionContext: As Seen by Org B The context received by Org B's extraction workflow, showing the full caller chain and governance configuration: ```json { "caller_chain": [ { "org_id": "org-c-uuid", "org_name": "Apex Consulting", "roles": ["caller"] }, { "org_id": "org-a-uuid", "org_name": "ResearchBot Inc", "roles": ["marketplace_provider"] }, { "org_id": "org-b-uuid", "org_name": "MedExtract Co", "roles": ["marketplace_provider"] } ], "project_id": "proj-clinical-review-2026", "user_id": "user-jane-doe-uuid", "budget": { "total": "8.00", "spent": "0.00", "reserved": "0.00" }, "trace_id": "trace-a1b2c3d4e5f6", "parent_span_id": "span-org-a-agent-7890", "span_id": "span-org-b-extraction-1234", "prompt_stack": [ { "source": "global", "role": "policy", "content": "You are operating within the AceTeam platform. Follow all data governance policies. Do not expose PHI to unauthorized parties.", "priority": 100 }, { "source": "org:org-a-uuid", "role": "persona", "content": "You are a clinical research assistant. Cite all sources. Flag uncertain findings.", "priority": 50 }, { "source": "agent:agent-extraction-uuid", "role": "task", "content": "Extract entities and relations from the provided clinical document. Use ICD-10 codes where applicable.", "priority": 10 } ], "governance_context": { "consent_endpoint": "https://consent.aceteam.internal/api/v1", "classification_policies": [ { "org_id": "org-b-uuid", "default_level": "CONFIDENTIAL", "phi_detection_enabled": true } ] }, "values": { "deployment_region": "us-east-1", "model_preference": "gpt-4o" } } ``` Note that Org B's budget shows `"total": "8.00"`, the sub-budget allocated by Org A. Org B cannot see Org C's original $50.00 budget. ### B.2 ExecutionEnvelope: Returned by Org A to Org C The complete envelope returned to Org C after the full execution tree completes: ```json { "result": { "@type": "type.googleapis.com/aep.core.v1.AgentResponse", "content": "Based on the clinical document analysis, the patient presents with Type 2 Diabetes Mellitus (ICD-10: E11.9) with peripheral neuropathy (G63). Key findings:\n\n1. HbA1c: 8.2% (above target) [ref:span-org-b-extraction-1234]\n2. Fasting glucose: 186 mg/dL [ref:span-org-b-extraction-1234]\n3. Nerve conduction velocity: reduced in bilateral lower extremities [ref:span-org-b-extraction-1234]\n\nRecommendation: Intensify glycemic control and refer to neurology.", "tool_calls_made": 3 }, "cost_tree": { "span_id": "span-org-a-agent-7890", "executor_org": "org-a-uuid", "caller_org": "org-c-uuid", "service_type": "agent", "service_id": "svc-research-agent", "compute_cost": "5.00", "value_added_fee": "5.00", "platform_fee": "0.75", "children": [ { "span_id": "span-org-b-extraction-1234", "executor_org": "org-b-uuid", "caller_org": "org-a-uuid", "service_type": "workflow", "service_id": "svc-clinical-extraction", "compute_cost": "3.00", "value_added_fee": "2.00", "platform_fee": "0.25", "children": [ { "span_id": "span-org-b-gliner-5678", "executor_org": "org-b-uuid", "caller_org": "org-b-uuid", "service_type": "tool", "service_id": "gliner2-entity-extraction", "compute_cost": "1.50", "value_added_fee": "0.00", "platform_fee": "0.00", "children": [] }, { "span_id": "span-org-b-llm-9012", "executor_org": "org-b-uuid", "caller_org": "org-b-uuid", "service_type": "tool", "service_id": "gpt-4o-structured-output", "compute_cost": "1.50", "value_added_fee": "0.00", "platform_fee": "0.00", "children": [] } ] } ] }, "spans": [ { "span_id": "span-org-a-agent-7890", "parent_span_id": "span-org-c-root-0001", "trace_id": "trace-a1b2c3d4e5f6", "executor_type": "agent", "executor_id": "agent-research-uuid", "org_id": "org-a-uuid", "start_time": "2026-02-04T10:00:00.000Z", "end_time": "2026-02-04T10:00:12.450Z", "status": "OK", "metadata": { "model": "gpt-4o", "prompt_tokens": 3200, "completion_tokens": 480 } }, { "span_id": "span-org-b-extraction-1234", "parent_span_id": "span-org-a-agent-7890", "trace_id": "trace-a1b2c3d4e5f6", "executor_type": "workflow", "executor_id": "workflow-clinical-extraction-uuid", "org_id": "org-b-uuid", "start_time": "2026-02-04T10:00:03.100Z", "end_time": "2026-02-04T10:00:08.750Z", "status": "OK", "metadata": { "nodes_executed": 4, "entities_extracted": 12 } } ], "citations": [ { "span_id": "span-org-b-extraction-1234", "source_org": "org-b-uuid", "source_type": "workflow", "source_id": "workflow-clinical-extraction-uuid", "content": "HbA1c: 8.2%, Fasting glucose: 186 mg/dL, Nerve conduction velocity: reduced bilateral LE", "confidence": 0.94, "classifications": [ { "level": "PHI", "regulations": ["HIPAA"], "data_subjects": ["patient:uuid-pat-5678"], "field_path": "result.entities[0].properties.hba1c", "source_span_id": "span-org-b-extraction-1234", "constraints": { "allow_redistribution": false, "requires_encryption": true, "retention_policy": "7y", "allowed_purposes": ["treatment", "billing"], "redaction_required": false, "audit_required": true } } ] } ], "errors": [] } ``` ### B.3 ConsentGrant: Patient Authorization ```json { "grant_id": "consent-grant-abc123", "data_subject": "patient:uuid-pat-5678", "granted_to": "org:org-a-uuid", "purposes": ["treatment", "billing"], "fields": ["*"], "expiry": "2027-02-04T00:00:00Z", "revocable": true, "granted_by": "patient:uuid-pat-5678", "granted_at": "2026-01-15T09:30:00Z" } ``` ### B.4 ConsentDecision: Boundary Check Result ```json { "allowed": true, "reason": "Active consent grant consent-grant-abc123 covers purpose 'treatment' for all fields", "constraints": ["audit_required", "no_redistribution"], "expires_at": "2027-02-04T00:00:00Z", "audit_record_id": "audit-log-xyz789" } ``` ### B.5 BudgetState: Mid-Execution Snapshot Budget state at Org A after two LLM calls, with one in-flight marketplace call to Org B: ```json { "total": "50.00", "spent": "12.35", "reserved": "8.00" } ``` Available: $50.00 - $12.35 - $8.00 = **$29.65** ### B.6 Job Queue Payload: With AEP Extension ```json { "version": "1.0", "type": "agent_response", "jobId": "job-uuid-1234", "rayId": "ray-SFO-2m5x7k-a1b2c3", "userId": "user-jane-doe-uuid", "organizationId": "org-c-uuid", "createdAt": "2026-02-04T10:00:00.000Z", "priority": 0, "maxAttempts": 3, "timeoutSeconds": 300, "execution_context": { "caller_chain": [ { "org_id": "org-c-uuid", "org_name": "Apex Consulting", "roles": ["caller"] } ], "project_id": "proj-clinical-review-2026", "budget": { "total": "50.00", "spent": "0.00", "reserved": "0.00" }, "trace_id": "trace-a1b2c3d4e5f6", "span_id": "span-org-c-root-0001", "parent_span_id": "", "prompt_stack": [ { "source": "global", "role": "policy", "content": "You are operating within the AceTeam platform.", "priority": 100 } ], "governance_context": { "consent_endpoint": "https://consent.aceteam.internal/api/v1", "classification_policies": [] } } } ``` ## Appendix C: Relationship to Existing Systems | Existing Component | AEP Equivalent | Notes | | ---------------------------------------------------------------------------- | ----------------------------------------- | ---------------------------------------------------------- | | `BaseJobPayload.rayId` | `ExecutionContext.trace_id` | `rayId` retained for backward compatibility | | `BaseJobPayload.userId` | `ExecutionContext.user_id` | Direct mapping | | `BaseJobPayload.organizationId` | `ExecutionContext.caller_chain[0].org_id` | Originating org | | `StreamEventType` (start, chunk, end, error, tool_call_start, tool_call_end) | Retained, plus 5 new AEP types | Backward compatible | | `TokenUsage` (prompt_tokens, completion_tokens, cost) | `CostNode` + `Span.metadata` | Cost tracked hierarchically; token counts in span metadata | | `JobStatus` (enqueued, claimed, processing, completed, failed) | Retained | No change to job lifecycle | | `requiredCapabilities` (tag-based routing) | Retained | Used for fabric mesh routing | | `stream:v1:{jobId}` (Redis Pub/Sub channel) | Retained | New event types added to same channel | | `model_usage_stats` (per-call cost tracking) | `CostNode` (hierarchical cost tree) | Flat tracking → recursive tree | | `pythonBackendFetch` (Next.js → Python) | Same pattern, AEP context in headers | REST transport preserved | ## Appendix D: Glossary | Term | Definition | | --------------------------- | ------------------------------------------------------------------------------------------------------------------------ | | **AEP** | Agentic Execution Protocol. The vendor-neutral wire protocol defined in this document. | | **Caller chain** | Ordered list of organizational identities representing the chain of callers from originator to current executor. | | **Citation** | A structured reference linking part of an executor's output to the source span, organization, and data that produced it. | | **Conformance level** | One of four tiers (0–3) declaring which AEP features an implementation supports. | | **Consent grant** | A record authorizing a specific organization to access classified data for specific purposes. | | **CostNode** | A recursive tree structure recording compute cost, value-added fee, and platform fee for each executor in the call tree. | | **Data classification** | A label on a data field specifying its security level, applicable regulations, and handling constraints. | | **Execution context** | The downward-flowing message carrying identity, budget, tracing, prompts, and governance policy. | | **Execution envelope** | The upward-flowing message carrying results, cost trees, spans, citations, and errors. | | **Executor** | Any component that implements `Execute(ExecutionContext) → (ExecutionEnvelope, error)`. | | **Fabric mesh** | A WireGuard-based VPN connecting organizational infrastructure for cross-boundary executor calls. | | **Org boundary** | The point where execution crosses from one organization's infrastructure to another's. | | **Pessimistic reservation** | Budget enforcement strategy: reserve worst-case cost before execution, settle with actual cost after. | | **Prompt layer** | A single layer in a composable prompt stack, with a source, role, content, and priority. | | **Span** | A structured record of a single unit of work, forming a tree via parent-child relationships. | | **Sub-budget** | A budget allocated from the caller's available budget to a callee at an organizational boundary. | --- ## AEP Safety Proxy Source: https://aceteam.ai/docs/safety-proxy A reverse proxy that intercepts every LLM call, detects PII and toxic content, enforces PASS/FLAG/BLOCK decisions, and tracks cost. Zero code changes. # AEP Safety Proxy The AEP safety proxy sits between your agent and the LLM API. It reads every request and every response. It can block in either direction: before the call reaches the API, or before the agent sees the response. ## Install and run ```bash pip install aceteam-aep[all] aceteam-aep proxy --port 8080 ``` Point any agent at the proxy: ```bash export OPENAI_BASE_URL=http://localhost:8080/v1 # Run your agent normally — every call goes through the proxy ``` The proxy serves a real-time dashboard at `http://localhost:8080/aep/`. ## What the proxy sees The proxy is a man-in-the-middle between your agent and the LLM API. It has full visibility into the wire traffic. | Direction | What's visible | | ------------ | ------------------------------------------------------------------------ | | **Request** | Every message sent to the LLM, including system prompts and tool results | | **Response** | The full LLM response: text, tool calls, token usage | | **Cost** | Per-call token counts and dollar cost (using model-specific pricing) | What the proxy **cannot** see (use [AEP headers](#aep-headers) or the Python SDK for these): - Agent actions between calls (file writes, code execution, browser actions) - Application context (who is calling, data classification, consent) ## Two layers Think WireGuard + Tailscale. WireGuard is a minimal wire protocol: fast, free, ubiquitous. Tailscale adds identity and management on top. **Layer 1: AEP Proxy (free)**. Wire-level reverse proxy. Zero code changes. Intercepts both directions. PII detection, toxicity classification, cost anomaly detection, PASS/FLAG/BLOCK enforcement. Real-time dashboard. Works with any language or framework that calls the OpenAI-compatible API. **Layer 2: AEP SDK + Headers (enterprise)**. Application-level protocol. Adds identity (`X-AEP-Entity`), governance (`X-AEP-Classification`, `X-AEP-Consent`), and provenance (citation chains). Injected via HTTP headers through the proxy or via the Python SDK. ## Built-in safety detectors The proxy ships with three detectors enabled by default: ### PII Detection Catches personally identifiable information using a local NER transformer model (`iiiorg/piiranha-v1-detect-personal-information`, ~110M parameters). Falls back to regex patterns if the model isn't available. **Detected entity types:** SSN, email, phone, credit card, IP address, person name, ID numbers. - Signal type: `pii` - Severity: `high` - Default action: **BLOCK** ### Content Safety Classifies text as safe/unsafe using a local toxicity classifier (`s-nlp/roberta_toxicity_classifier`, ~125M parameters). Checks both input and output text. Threshold configurable (default: 0.7). - Signal type: `content_safety` - Severity: `high` (score > 0.9) or `medium` - Default action: **BLOCK** (high) or **FLAG** (medium) ### Cost Anomaly Flags calls whose cost exceeds 5x the session average (configurable). Only activates after a baseline of 3 calls. Catches runaway loops before they become surprise bills. - Signal type: `cost_anomaly` - Severity: `medium` - Default action: **FLAG** All detectors run locally on CPU. No API calls, no data sent to third parties. ## Enforcement Every call produces an `EnforcementDecision`: **PASS**, **FLAG**, or **BLOCK**. The default `EnforcementPolicy`: | Severity | Action | | -------- | -------------------------------------- | | `high` | BLOCK: request or response is rejected | | `medium` | FLAG: logged and surfaced, not blocked | | `low` | PASS | The policy is configurable: ```python from aceteam_aep.enforcement import EnforcementPolicy # Block on high severity, flag on medium, exempt cost_anomaly from enforcement policy = EnforcementPolicy( block_on=frozenset({"high"}), flag_on=frozenset({"medium"}), allow_types=frozenset({"cost_anomaly"}), ) ``` ## AEP headers The proxy reads and strips `X-AEP-*` headers from requests before forwarding to the LLM API. This is the bridge between Layer 1 and Layer 2. | Header | Purpose | Example | | ---------------------- | -------------------------------- | ------------------------------ | | `X-AEP-Entity` | Identity: who is making the call | `org:acme-corp` | | `X-AEP-Classification` | Data sensitivity level | `confidential` | | `X-AEP-Consent` | Data governance consent | `gdpr:granted` | | `X-AEP-Budget` | Spend limit for this call | `0.50` | | `X-AEP-Sources` | Declared data sources | `q3-report.pdf,crm-export.csv` | Response headers added by the proxy: | Header | Purpose | | ------------------- | ---------------------------------------------- | | `X-AEP-Cost` | Dollar cost of this call | | `X-AEP-Enforcement` | Enforcement decision (pass/flag/block) | | `X-AEP-Call-ID` | Unique call identifier | | `X-AEP-Flag-Reason` | Reason string (only present on FLAG decisions) | ## Streaming The proxy supports SSE streaming (`stream: true`). Chunks are passed through to the agent immediately. Safety detectors run on the complete response after the stream finishes. If the post-stream check triggers a BLOCK, the proxy appends an `aep_safety_block` SSE event. ## Framework compatibility The proxy works with anything that calls the OpenAI-compatible API: - **Python SDK** (`openai`, `anthropic`): Set `base_url` or `OPENAI_BASE_URL` env var - **LangChain**: Configure `openai_api_base` on the LLM - **CrewAI**: Uses OpenAI client under the hood. Set `OPENAI_BASE_URL` - **DSPy**: Uses LiteLLM. Set `OPENAI_API_BASE` or `api_base` on `dspy.LM` - **OpenClaw**: Custom model provider config pointing at the proxy - **curl**: `curl http://localhost:8080/v1/chat/completions -H "Authorization: Bearer $OPENAI_API_KEY" ...` ## Hosted alternative Don't want to run the proxy yourself? The [AceTeam Gateway](/docs/gateway) is a hosted version of the same safety pipeline with additional features: multi-provider model routing, org-level cost tracking, and credit-based billing. Switch by changing one environment variable. ## Next steps - [Gateway API](/docs/gateway): hosted OpenAI-compatible proxy with model routing - [Safety Detectors](/docs/safety-detectors): write custom detectors, detector architecture - [AEP Protocol overview](/docs/aep-overview): the full accountability protocol - [Quickstart](https://github.com/aceteam-ai/aep-quickstart): clone and run in 5 minutes --- ## Safety Detectors Source: https://aceteam.ai/docs/safety-detectors How AEP's pluggable detector architecture works. Built-in detectors for PII, toxicity, and cost anomalies, plus how to write your own. # Safety Detectors AEP uses a pluggable detector architecture. Each detector is a Python class that implements the `SafetyDetector` protocol: a `name` attribute and a `check()` method. The `DetectorRegistry` runs all registered detectors on every call and aggregates their signals. ## SafetyDetector protocol ```python from aceteam_aep.safety import SafetySignal class MyDetector: name: str = "my_detector" def check( self, *, input_text: str, output_text: str, call_id: str, **kwargs: object, ) -> list[SafetySignal]: ... ``` That's the entire interface. Any class with `name` and `check()` works. No base class inheritance required: it's a Python protocol (structural typing). ## SafetySignal Each detector returns a list of `SafetySignal` objects: ```python from aceteam_aep.safety import SafetySignal signal = SafetySignal( signal_type="my_check", # Category of the signal severity="high", # "low", "medium", or "high" call_id="call_123", # Which call triggered this detail="Found issue X", # Human-readable description ) # detector and timestamp are auto-populated ``` | Field | Type | Description | | ------------- | ----- | -------------------------------------------------------------------- | | `signal_type` | `str` | Category: `"pii"`, `"content_safety"`, `"cost_anomaly"`, or your own | | `severity` | `str` | `"low"`, `"medium"`, or `"high"` (maps to enforcement actions) | | `call_id` | `str` | Identifies which API call triggered the signal | | `detail` | `str` | Human-readable description of the issue | | `detector` | `str` | Auto-populated from the detector's `name` | | `timestamp` | `str` | Auto-generated ISO 8601 timestamp | ## Built-in detectors ### PiiDetector Detects personally identifiable information in LLM output text. **Primary method:** Local NER transformer model (`iiiorg/piiranha-v1-detect-personal-information`, ~110M parameters). Runs on CPU. Lazy-loaded on first call. **Fallback:** Regex patterns for SSN, credit cards, email, and phone numbers. Used when `transformers` is not installed. **Detected entities:** SSN, EMAIL, PHONE, CREDIT_CARD, IP_ADDRESS, PERSON, ID_NUM. **Configuration:** ```python from aceteam_aep.safety.pii import PiiDetector detector = PiiDetector( model_name="iiiorg/piiranha-v1-detect-personal-information", # default ) ``` Text is truncated to 2048 characters to prevent OOM. Results are deduplicated by entity type. ### ContentSafetyDetector Classifies text as safe or unsafe using a local toxicity model. **Model:** `s-nlp/roberta_toxicity_classifier` (~125M parameters). Runs on CPU. **Configuration:** ```python from aceteam_aep.safety.content import ContentSafetyDetector detector = ContentSafetyDetector( model_name="s-nlp/roberta_toxicity_classifier", # default threshold=0.7, # classification score threshold ) ``` Checks both input and output text (truncated to 512 characters each). Severity is `high` if score > 0.9, otherwise `medium`. ### CostAnomalyDetector Statistical detector that flags calls whose cost exceeds a multiple of the session average. **Configuration:** ```python from aceteam_aep.safety.cost_anomaly import CostAnomalyDetector detector = CostAnomalyDetector( min_calls=3, # baseline calls before detection activates multiplier=5, # threshold: 5x session average ) ``` Passes `call_cost` via kwargs. Only activates after `min_calls` calls have been recorded (first fires on call `min_calls + 1`). ## Writing a custom detector ```python from aceteam_aep.safety import SafetySignal class ProfanityDetector: name = "profanity" BLOCKED_WORDS = {"badword1", "badword2"} def check( self, *, input_text: str, output_text: str, call_id: str, **kwargs: object, ) -> list[SafetySignal]: signals = [] for word in self.BLOCKED_WORDS: if word in output_text.lower(): signals.append( SafetySignal( signal_type="profanity", severity="medium", call_id=call_id, detail=f"Blocked word detected in output", ) ) return signals ``` Register it: ```python from aceteam_aep import wrap import openai client = wrap( openai.OpenAI(), detectors=[ProfanityDetector()], ) ``` Or with the proxy: ```python from aceteam_aep.proxy.app import create_proxy_app app = create_proxy_app(detectors=[ProfanityDetector()]) ``` ## DetectorRegistry The `DetectorRegistry` manages detector execution. Key behavior: - **Fault isolation:** If one detector throws, others still run. Failures are logged, never propagated to the caller. - **Auto-naming:** The `detector` field on each `SafetySignal` is automatically set to the detector's `name`. - **Aggregation:** All signals from all detectors are collected into a single list. ```python from aceteam_aep.safety import DetectorRegistry registry = DetectorRegistry() registry.add(PiiDetector()) registry.add(ContentSafetyDetector()) registry.add(ProfanityDetector()) signals = registry.run_all( input_text="Hello", output_text="Response with SSN 123-45-6789", call_id="call_001", ) ``` ## Enforcement mapping Signals flow into the `EnforcementPolicy` which maps severities to actions: | Severity | Default Action | Meaning | | -------- | -------------- | ------------------------------------------ | | `high` | **BLOCK** | Request or response is rejected | | `medium` | **FLAG** | Logged, surfaced in dashboard, not blocked | | `low` | **PASS** | Recorded for audit, no action | See [Safety Proxy](/docs/safety-proxy) for enforcement policy configuration. --- ## Trust Engine Source: https://aceteam.ai/docs/safety-trust-engine AceTeam's ensemble-of-judges approach to AI safety evaluation. Calibrated confidence scores from diverse judge models, validated by academic research. # Trust Engine The Trust Engine is AceTeam's approach to evaluating whether AI agent outputs are safe, accurate, and compliant. Instead of relying on a single classifier, it uses an ensemble of independent judge models that produce calibrated confidence scores. ## The problem A single safety classifier has a single failure mode. If the model was trained on data that doesn't cover your domain, it will miss things. If it has a systematic bias, every call inherits that bias. Cost tracking and PII detection (the [built-in detectors](/docs/safety-detectors)) catch the obvious cases. The Trust Engine handles the hard cases: nuanced content policy, domain-specific safety requirements, and outputs where "safe" depends on context. ## How it works The Trust Engine evaluates agent outputs using a **diverse panel of independent judge models**. Each judge independently assesses whether an action is safe, and their calibrated agreement produces a confidence score. ``` Agent output ├── Judge 1 (safety classifier) → safe (0.82) ├── Judge 2 (policy compliance) → safe (0.91) └── Judge 3 (domain expert) → flagged (0.34) ↓ Ensemble score: 0.69 Decision: FLAG (below threshold) ``` When multiple judges with uncorrelated error patterns all agree an output is correct, that's a stronger signal than any single model's assessment. When they disagree, the output gets flagged for review. ## Research validation The ensemble approach is backed by two lines of evidence: ### AceTeam + University of Waterloo / Vector Institute Active research collaboration with Prof. Pascal Poupart. Current results on academic benchmarks show the ensemble approach improving with diversity: | Judges | AUC (discrimination) | Calibration Error | | ------ | -------------------- | ----------------- | | 1 | 0.58 | 0.70 | | 2 | 0.63 | 0.55 | | 3 | 0.68 | 0.45 | More diverse judges = better calibration. The trajectory is clear. ### Independent validation (Lin et al., March 2026) [Lin et al.](https://arxiv.org/abs/2511.06396v3), from the Beijing Institute of AI Safety and Governance, showed that structured multi-agent debate with small models (14B parameters) achieves **97% of GPT-4o's accuracy** on safety evaluation at **46% of the cost**. Three debate rounds is optimal. This means the Trust Engine can run on cheap inference, critical for a system that evaluates every agent call. ## Architecture The Trust Engine sits above the built-in detectors in the safety stack: ``` Layer 3: Trust Engine (ensemble judges, calibrated confidence) Layer 2: ML Detectors (PII, toxicity, content safety) Layer 1: Pattern Detectors (regex, cost anomaly) ``` Each layer adds cost and accuracy. Layer 1 is near-zero cost. Layer 2 runs small models (~110-125M params) per call. Layer 3 runs 3-5 judge models (each ~7B params) per call. ## Business model alignment The compute cost of safety evaluation IS the product: | Tier | What runs | Compute cost | Price | | -------------- | ----------------------------------------------------- | ----------------------------- | ------ | | **Free** | Pattern detectors, cost tracking, dashboard | Near-zero | $0 | | **Pro** | ML-based PII, toxicity, content safety | ~3 small models/call | $20/mo | | **Enterprise** | Trust Engine (ensemble judges, calibrated confidence) | 3-5 x 7B model inference/call | Custom | Free observation gets developers in the door. Active safety classification costs tokens. That cost gap is the revenue. ## Detector integration The Trust Engine implements the same `SafetyDetector` protocol as all other detectors. It plugs into the existing `DetectorRegistry`: no special infrastructure required. ```python # Future API (Trust Engine is in active development) from aceteam_aep.trust_engine import TrustEngine engine = TrustEngine( judges=["safety-classifier-7b", "policy-compliance-7b", "domain-expert-7b"], threshold=0.7, ) # Registers as a standard detector client = wrap(openai.OpenAI(), detectors=[engine]) ``` ## Status The Trust Engine is in active development. The built-in detectors (PII, toxicity, cost anomaly) are production-ready in [aceteam-aep v0.4.0](https://pypi.org/project/aceteam-aep/). The ensemble judge system is being built in collaboration with UWaterloo/Vector Institute. ## Next steps - [Safety Proxy](/docs/safety-proxy): get started with the free safety proxy - [Safety Detectors](/docs/safety-detectors): write custom detectors - [AEP Protocol](/docs/aep-overview): the full accountability protocol --- ## Hosted Instances Source: https://aceteam.ai/docs/instances-overview Deploy AI agents as hosted containers with automatic safety enforcement. No infrastructure to manage: create an instance, pick a template, and your agent runs 24/7 with cost tracking and audit trails. # Hosted Instances Hosted instances let you run AI agents as managed containers on AceTeam's infrastructure. Every instance gets automatic safety enforcement: your agent's LLM calls are routed through the AEP Safety Proxy with cost tracking, governance, and audit trails. You bring the agent. We handle hosting, safety, and billing. ## How it works When you create an instance, three things happen: 1. **A container starts** with your chosen agent image (from a template or your own Docker image) 2. **A safety session is created**: an AEP Gateway session with your chosen policy and budget 3. **The container is wired up**: `OPENAI_BASE_URL` and `OPENAI_API_KEY` are injected so all LLM calls route through safety Your agent doesn't need any code changes. It thinks it's talking to OpenAI. Behind the scenes, every call goes through safety detection, cost tracking, and governance enforcement. ``` Your Agent (in container) │ │ OPENAI_BASE_URL → AEP Gateway │ OPENAI_API_KEY → act_xxxx ▼ AEP Safety Proxy ├── Safety detection (PII, toxicity, agent threats) ├── Cost tracking (per-call, per-session budgets) ├── Governance (classification, consent, audit trail) ▼ LLM API (OpenAI, Anthropic, etc.) ``` ## What you get | Feature | Description | | ------------------------ | --------------------------------------------------- | | **Container hosting** | Your agent runs 24/7 on managed infrastructure | | **Safety enforcement** | Every LLM call checked by safety detectors | | **Cost tracking** | Per-instance budgets with hard caps | | **Audit trail** | Every call logged with safety signals and decisions | | **Lifecycle management** | Start, stop, restart from the dashboard | | **Container logs** | View your agent's output in real-time | ## Templates Templates are pre-configured agent images ready to deploy. Select a template when creating an instance: it sets the Docker image, default configuration, and resource limits. | Template | Description | Image | | -------------- | ----------------------------------------- | ------------------------------------------ | | SafeClaw Agent | AI agent with SafeClaw safety enforcement | `ghcr.io/aceteam-ai/safeclaw-agent:latest` | More templates are coming. You can also deploy any Docker image using the "Custom Image" option. ## Custom images Bring your own Docker image. Any agent framework that supports the OpenAI-compatible API works automatically: - OpenClaw / SafeClaw agents - LangChain agents - Any agent using the `openai` Python or Node SDK - Custom agents calling `OPENAI_BASE_URL` Your container receives these environment variables automatically: | Variable | Value | Purpose | | ----------------- | ----------------------------------- | -------------------------------------- | | `OPENAI_BASE_URL` | `https://aceteam.ai/api/gateway/v1` | Routes LLM calls through safety | | `OPENAI_API_KEY` | `act_xxxx` | Authenticates with your safety session | | `AEP_GATEWAY_URL` | Same as `OPENAI_BASE_URL` | Explicit AEP reference | | `AEP_INSTANCE_ID` | Instance UUID | Links calls to this instance | You can add your own environment variables when creating the instance. ## Safety policies Each instance has a safety policy that controls how safety signals are handled: | Policy | Behavior | | -------------- | ------------------------------------------------------------ | | **Default** | Block high-severity threats, flag medium-severity for review | | **Strict** | Block both high and medium severity signals | | **Permissive** | Flag everything, block nothing (monitoring only) | ## Budget caps Set a maximum spend per instance in USD. When the budget is reached, further LLM calls are blocked. Monitor spend in real-time from the instance detail page or the gateway dashboard. ## Instance states | State | Meaning | | -------- | ------------------------------------------------- | | Created | Instance configured but container not yet started | | Starting | Container is being deployed | | Running | Container is active and processing requests | | Stopping | Container is shutting down | | Stopped | Container stopped, configuration preserved | | Error | Container failed to start or crashed | You can restart a stopped or errored instance at any time. Deleting an instance stops the container and removes the configuration. ## Limits Each instance runs with configurable resource limits: - **Memory:** 128 MB to 4 GB (default: 1 GB) - **CPU:** 100m to 4000m (default: 500m) ## What's next - **Persistent memory**: agents that remember across conversations - **Scheduled tasks**: cron jobs with full agent context - **Multi-channel**: connect WhatsApp, Telegram, Discord to your instance - **More templates**: Hermes, custom framework adapters --- ## Quickstart: Deploy Your First Instance Source: https://aceteam.ai/docs/instances-quickstart Create a hosted AI agent instance in under a minute. Pick a template, set a safety policy, and your agent is live with full safety enforcement. # Deploy Your First Instance This guide walks you through creating and managing a hosted agent instance. By the end, you'll have an AI agent running 24/7 with automatic safety enforcement. **Time required:** Under 1 minute. **Prerequisites:** An AceTeam account with credits. [Sign up](https://aceteam.ai) if you haven't: you get $1 free credit on signup. ## 1. Create an instance Go to [aceteam.ai/instances](https://aceteam.ai/instances) and click **New Instance**. **Pick a template.** Select "SafeClaw Agent" for a pre-configured agent with safety enforcement. Or select "Custom Image" to deploy your own Docker image. **Name it.** Give your instance a descriptive name (e.g., "Research Assistant" or "Customer Support Bot"). **Set a safety policy.** - **Default**: blocks dangerous requests, flags suspicious ones for review - **Strict**: blocks anything that triggers a safety signal - **Permissive**: monitors everything but blocks nothing (useful for testing) **Set a budget cap.** Maximum USD spend for this instance. Start with $2 for testing. Click **Create Instance**. ## 2. Save the gateway API key After creation, you'll see a gateway API key. This key is injected into your container automatically: you don't need to configure it manually. But save a copy for debugging: - Click **Copy** to copy to clipboard - Click **Download** to save as a text file This key is shown only once. ## 3. Start the instance From the instance card on the dashboard, click **Start**. The state indicator changes: - Gray dot = Created (not started) - Amber pulsing dot = Starting (container deploying) - Green dot = Running (agent is live) - Red dot = Error (check logs) Starting typically takes 15-30 seconds. ## 4. Monitor your instance Click the instance name to open the detail page. Here you can: - **View logs**: see your agent's console output in real-time - **View gateway dashboard**: see safety signals, cost tracking, and call history - **Stop or restart**: control the instance lifecycle ## 5. Connect channels (coming soon) Soon you'll be able to bind WhatsApp, Telegram, or Discord channels to your instance. Messages from those channels will be routed to your agent with full safety enforcement. ## Using your own Docker image If you selected "Custom Image" instead of a template: 1. Your image must be publicly pullable (or hosted on `ghcr.io/aceteam-ai/`) 2. Your agent must read `OPENAI_BASE_URL` and `OPENAI_API_KEY` from the environment 3. That's it: the AEP Gateway handles the rest Example with a LangChain agent: ```python from langchain_openai import ChatOpenAI import os # These are injected automatically by the instance llm = ChatOpenAI( base_url=os.environ["OPENAI_BASE_URL"], api_key=os.environ["OPENAI_API_KEY"], ) ``` Your agent doesn't need to know about AEP. It calls the OpenAI API as normal, and every call is safety-checked transparently. ## Next steps - [Hosted Instances overview](/docs/instances-overview): full reference for templates, policies, states, and limits - [AEP Safety Proxy](/docs/safety-proxy): understand the safety layer your instance uses - [Safety Detectors](/docs/safety-detectors): what the safety system checks for --- ## AceTeam Executive Overview Source: https://aceteam.ai/docs/executive-overview A business-oriented introduction to AceTeam: AI infrastructure you own and control. Artificial intelligence is reshaping every industry, from healthcare and finance to education and consulting. But organizations adopting AI face a difficult choice: hand your data to a cloud provider for convenience, or build everything yourself for control. AceTeam eliminates that trade-off. We give you the convenience of a managed AI platform with the control of running on your own infrastructure. ## The Challenge Organizations trying to adopt AI today run into three problems that slow them down or stop them entirely. **Your data leaves your hands.** When you use a cloud AI service, every question you ask and every document you process travels to someone else's servers. You do not control where that data is stored, who can access it, or what jurisdiction it falls under. For organizations handling sensitive client information, patient records, or proprietary research, this is not just uncomfortable: it can be a regulatory violation. **Costs are unpredictable and grow fast.** Cloud AI services charge per use: per question asked, per document processed, per image generated. A successful pilot becomes an expensive production deployment. A spike in usage becomes a budget emergency. Organizations that already own powerful computing hardware end up paying cloud fees on top of equipment they have already purchased. **Switching providers means starting over.** Once you build your AI tools on one provider's platform, moving to another is painful and expensive. Your agents, workflows, and integrations are tied to that provider's specific way of doing things. This gives the provider leverage over your pricing, your roadmap, and your options. ## What AceTeam Does AceTeam is an AI platform that runs on infrastructure you control. It provides three core capabilities. **Build intelligent agents.** Create virtual assistants that understand your organization's knowledge, speak in your brand's voice, connect to your tools, and interact with users through chat, voice, or phone. No coding required: configure agents through a visual interface and deploy them as web widgets, internal tools, or phone lines. **Automate complex workflows.** Design multi-step business processes using a visual drag-and-drop editor with over 30 building blocks. Combine AI reasoning with data lookups, approvals, conditional logic, and external system calls into repeatable, auditable pipelines. Build once, run thousands of times. **Run on your own hardware.** Install a single lightweight program on your servers, and they become part of your managed AI infrastructure. Your data stays on your machines, in your buildings, under your control. Add more hardware when you need more capacity. Mix your own servers with cloud resources when it makes sense. ## Who It's For **Enterprise teams** with GPU servers scattered across departments can unify them into a single AI platform. Engineering, research, and operations share one pool of computing resources with proper access controls and usage tracking, instead of each team managing its own fragmented setup. **Consulting firms** managing AI solutions for multiple clients get full data isolation between engagements. Build agents and workflows in your own environment, deploy them to client environments, and guarantee that no client's data ever touches another client's infrastructure. **Regulated industries** (healthcare, finance, government, legal) can adopt AI without compromising compliance. When your data never leaves your premises, meeting requirements like HIPAA, GDPR, and PIPEDA becomes straightforward rather than a barrier to adoption. **Universities and training programs** preparing the next generation of AI professionals can provide hands-on experience with production-grade tools. Our partnership with San José State University's AI operator certification program demonstrates how AceTeam makes enterprise AI accessible to learners through guided, gamified training environments. **Infrastructure partners** (hardware vendors, data center operators, managed service providers) can pair their physical infrastructure with AceTeam's software to offer customers a complete, ready-to-use AI platform rather than raw computing power. ## Why It Matters Now Three forces are converging to make sovereign AI infrastructure urgent rather than optional. **Regulation is accelerating.** Data residency and AI governance requirements are expanding worldwide. The EU AI Act, evolving HIPAA guidance, Canadian privacy reforms, and sector-specific rules are raising the bar for where and how AI can process sensitive data. Organizations that depend entirely on third-party cloud AI face growing compliance exposure. **Cloud AI costs are compounding.** As organizations move from pilot projects to production workloads, cloud AI spending is growing faster than budgets can absorb. The most valuable AI applications (agents that reason across large knowledge bases, workflows that chain multiple AI steps together) are also the most expensive to run on per-use pricing. **The talent gap is real.** There are not enough AI engineers to serve every organization that needs AI capabilities. AceTeam bridges this gap by making it possible for business operators, consultants, and domain experts to build and deploy AI solutions through visual tools, without writing code. The people who best understand the problems are empowered to build the solutions. ## The Economics Running AI on your own hardware changes the cost equation fundamentally. Cloud AI providers charge premium per-use fees that reflect their infrastructure costs, their margins, and the convenience they provide. When you run the same workloads on hardware you already own, those per-use fees disappear. Organizations that switch to self-hosted AI through AceTeam typically see cost reductions of 80 to 97 percent compared to equivalent cloud API pricing. The savings are largest for high-volume, high-value workloads: exactly the use cases that matter most. AceTeam charges a platform fee of 10 to 15 percent of compute value for orchestration, management, and platform services. Compare that to the 70 to 90 percent gross margins that major cloud providers build into their AI pricing. You pay for the software that makes your hardware useful, not for renting someone else's hardware at a markup. For organizations that already own computing hardware, the math is especially clear. That equipment represents capital you have already invested. Running AI workloads on it through AceTeam turns idle infrastructure into productive capacity at near-zero marginal cost. ## No Vendor Lock-In AceTeam supports multiple AI providers simultaneously, including OpenAI, Anthropic, Google, and open-source models you host yourself. You can use cloud AI for some tasks and local models for others, switching between providers per agent or per workflow step. This means you are never dependent on a single AI provider's pricing, availability, or product decisions. If a better model becomes available from a new provider, you can adopt it without rebuilding your agents or workflows. Your AI strategy stays yours to direct. ## Open Source Foundation Key components of the AceTeam platform are released as open source under the MIT license. You can inspect, audit, and verify the software that runs on your infrastructure. This is not a black box: it is a platform built on transparency and trust. Open-source components include the workflow execution engine, the node agent that runs on your hardware, the command-line interface for local AI workflows, and the Python library of workflow building blocks. Organizations that need to understand exactly what software is running in their environment can read every line of code. ## What's Next AceTeam's vision extends beyond a single organization's infrastructure. We are building toward a future where organizations can share computing capacity with each other through a secure marketplace, turning surplus hardware into revenue and giving every organization access to the AI computing power they need, when they need it, without dependence on centralized cloud providers. In the near term, we are focused on smarter workload routing that automatically optimizes for cost and performance, support for large-scale deployments spanning hundreds of machines, and the marketplace infrastructure that will enable organizations to monetize their idle hardware. The goal is an AI economy where computing power is abundant, distributed, and owned by the organizations that use it. ## Learn More - **[Technical Whitepaper](/docs/whitepaper)**: Deep dive into AceTeam's architecture, components, and technical design - **[Platform Overview](/docs/platform-overview)**: Explore the full feature set of the AceTeam platform - **[Account Setup](/docs/account-setup)**: Create your account and start building - **[Fabric Overview](/docs/fabric-overview)**: Learn about the Sovereign Compute Fabric in more detail --- ## Sovereign AI Compute Platform Source: https://aceteam.ai/docs/whitepaper Technical whitepaper: how AceTeam turns distributed hardware into a managed AI factory. > **Looking for a non-technical overview?** Read the [Executive Overview](/docs/executive-overview) for a business-oriented introduction to AceTeam. AceTeam is a platform that lets organizations run AI workloads on their own hardware with the convenience of a managed cloud service. This document explains the problem we solve, how our architecture works, and why it matters for teams that need control over their AI infrastructure. ## The Problem The current landscape of enterprise AI is defined by a fundamental trade-off: you can have convenience or you can have control, but rarely both. ### Vendor Lock-In Most organizations building AI-powered products depend entirely on a handful of cloud providers -- OpenAI, AWS Bedrock, Azure OpenAI, Google Cloud Vertex. These platforms offer polished APIs and fast onboarding, but they come with deep structural dependencies. Switching providers means rewriting integrations, revalidating outputs, and renegotiating contracts. Once your agents, workflows, and data pipelines are built on a single provider's API, migrating away is expensive and slow. This lock-in extends beyond the inference layer. Cloud AI platforms bundle storage, vector databases, fine-tuning pipelines, and monitoring into proprietary ecosystems. Each additional service you adopt increases the cost of exit. ### Data Sovereignty Regulatory frameworks around the world increasingly restrict where data can be processed. GDPR in Europe, PIPEDA in Canada, HIPAA in the United States, and sector-specific rules in finance and defense all impose constraints on data residency and processing location. When you send a prompt to a cloud inference API, your data traverses networks and lands on hardware you do not control, in jurisdictions you may not have vetted. For organizations in healthcare, government, financial services, or legal practice, this is not an abstract risk. It is a compliance requirement that determines whether AI adoption is even possible. ### Cost Unpredictability Cloud AI pricing is consumption-based: per token, per API call, per minute of audio processed. This model makes costs difficult to forecast and impossible to cap. A successful product launch or a spike in usage can send bills into unexpected territory. Per-token pricing also penalizes the most valuable AI use cases -- long-context reasoning, multi-step agent workflows, and retrieval-augmented generation -- because they consume the most tokens. Organizations that already own capable hardware (GPU workstations, on-premise servers, co-located racks) are paying cloud API fees on top of capital they have already invested. ### The Missing Middle No existing platform bridges the gap between raw infrastructure and managed AI services. You can rent GPUs from a cloud provider and manage everything yourself -- model serving, load balancing, job scheduling, networking, monitoring -- or you can use a fully managed API and accept the trade-offs above. There is no option that lets you point a platform at your own hardware and get a managed AI experience on top of it. AceTeam fills that gap. ## Our Approach The platform is built on three architectural principles. **Separation of orchestration and compute.** The platform layer (agent builder, workflow engine, marketplace, billing) runs independently of the hardware that executes AI workloads. Orchestration logic does not need to know where a GPU is located. Compute nodes do not need to understand billing or user management. This separation means you can add, remove, or replace hardware without touching the application layer. **Queue-based workload distribution.** Every AI job (inference, embedding, RAG search, workflow execution) enters a durable job queue. Worker nodes pull jobs from the queue based on their capabilities. This pattern provides natural load balancing, backpressure handling, and fault tolerance. If a node goes offline, its pending jobs are picked up by another node. **Secure mesh networking.** All communication between the platform and your hardware travels over an encrypted WireGuard mesh network. Nodes connect to each other without opening firewall ports, configuring NAT rules, or exposing services to the public internet. The architecture is organized into four layers: ``` CONSUMERS Web App, Mobile App, API, CLI PLATFORM Marketplace, Billing, Agent Builder, Workflow Engine ORCHESTRATION Job Queue, Routing, Service Discovery MESH NETWORK Nexus VPN, Citadel Agents, Node Registry INFRASTRUCTURE Data Centers, GPUs, On-Premise, Edge, Cloud VMs ``` Consumers interact with the platform layer through standard web and API interfaces. The platform translates user actions into jobs. The orchestration layer routes jobs to the right nodes. The mesh network layer handles secure connectivity. The infrastructure layer is whatever hardware you provide -- your own servers, rented GPUs, edge devices, or a mix of all three. ## Key Components ### Nexus -- Mesh Coordinator Nexus is the coordination server for the encrypted mesh network. Built on Headscale (an open-source implementation of the Tailscale control plane), Nexus manages node registration, key exchange, and network policy. Every node in the network communicates over WireGuard tunnels, providing end-to-end encryption without requiring you to manage certificates or key rotation. Nexus provides zero-configuration networking. Nodes behind firewalls, NAT gateways, or carrier-grade NAT can join the mesh without port forwarding. The coordination server facilitates NAT traversal so that nodes discover each other automatically. Organization isolation is enforced at the network level. Each organization receives its own Headscale user scope. Nodes registered by one organization are invisible to every other organization. Network policy, preauth keys, and device lists are all scoped per organization, ensuring that multi-tenant deployments maintain strict boundaries. ### Citadel -- Node Agent Citadel is a lightweight Go binary that runs on each compute node. It is the only software you install on your hardware. Citadel handles four responsibilities: **Device registration.** Citadel uses the OAuth 2.0 device authorization flow to register nodes. You run a single command, visit a URL to approve the device, and the node joins your organization's mesh network. No manual key management or configuration file editing is required. **Status heartbeats.** Citadel periodically reports node health -- CPU and GPU utilization, available memory, loaded models, and network status -- back to the platform. The orchestration layer uses this information for routing decisions. **Workload execution.** When a job is routed to a node, Citadel receives the job specification and manages the execution lifecycle: pulling model weights if needed, running inference, streaming results back to the platform, and reporting completion or failure. **Capability advertisement.** Each node declares what it can do using structured capability tags (model availability, GPU type, memory, region, compliance certifications). The orchestration layer matches jobs to nodes based on these tags. ### Job Queue The job queue is the central nervous system of the platform. Built on Redis Streams, it provides durable, ordered, exactly-once delivery of jobs to worker nodes. **Consumer groups** allow multiple nodes to share a workload. Each job is delivered to exactly one consumer in the group, providing natural load distribution without a separate load balancer. **Dead letter queue (DLQ) with retry.** Failed jobs are moved to a dead letter queue after a configurable number of attempts. Operators can inspect, retry, or discard failed jobs. Transient failures (network timeouts, temporary resource exhaustion) are retried automatically. **Backpressure.** When all consumers are busy, the queue buffers incoming jobs. Producers receive backpressure signals so the platform can display accurate wait times and make throttling decisions. **Streaming results.** Long-running jobs (agent conversations, workflow executions) stream partial results back to the platform in real time using Redis Pub/Sub. The platform layer converts these into Server-Sent Events (SSE) for the frontend, so users see tokens appear as they are generated rather than waiting for the full response. ## Resource Routing Not all AI resources are interchangeable. The platform uses two routing strategies depending on the nature of the resource. ### Fungible Resources -- Queue-Based Routing LLM inference, embedding generation, and image generation are fungible workloads. Any node with the right model and sufficient capacity can handle the request. These jobs enter the queue and are picked up by the first available capable node. For example, if three nodes in your mesh advertise `llm:llama3` capability, an inference job for Llama 3 will be routed to whichever node is idle first. ### Specific Resources -- Direct Fabric Calls Some resources are not interchangeable. A database containing your proprietary data exists on a specific machine. A fine-tuned model checkpoint lives on a particular GPU node. Hardware with specific compliance certifications (SOC 2, HIPAA) may be required for certain workloads. For these cases, the platform routes requests directly to the target node over the mesh VPN. The job specification includes the target node identifier, and the orchestration layer delivers the request over the encrypted mesh without going through the general queue. ### Capability Tags Nodes advertise their capabilities using structured tags. The routing layer uses these tags to match jobs with appropriate hardware: - **Model tags**: `llm:llama3`, `llm:mixtral`, `embedding:bge-large` - **Hardware tags**: `gpu:a100`, `gpu:b300`, `ram:256gb` - **Location tags**: `region:us-east`, `region:eu-west`, `datacenter:dc-1` - **Compliance tags**: `compliance:soc2`, `compliance:hipaa`, `compliance:gdpr` Jobs specify required tags, and the routing layer ensures they land on nodes that satisfy all requirements. This allows organizations to express complex placement policies (for example, "run this medical inference job on a HIPAA-compliant node in the US-East region with at least one A100 GPU") without writing custom routing logic. ## AI Platform The Sovereign Compute Fabric is the infrastructure layer. On top of it, AceTeam provides a full AI application platform. ### Multi-Provider AI The platform supports multiple AI providers simultaneously: OpenAI, Anthropic Claude, Google Gemini, Deepseek, and local models via Ollama. You can use cloud APIs for some workloads and local inference for others, switching between providers per agent or per workflow step. This eliminates single-provider lock-in at the application layer. ### Agent Platform The agent builder lets you create AI agents with system prompts, knowledge bases (via retrieval-augmented generation), tool integrations (via the Model Context Protocol), real-time voice interaction (via WebRTC), and telephone connectivity (via Twilio). Agents can be deployed as web widgets, API endpoints, phone numbers, or internal tools. Knowledge bases are built using a RAG pipeline that ingests documents, chunks them, generates embeddings, and stores them in a vector database. At query time, relevant chunks are retrieved and included in the agent's context. The entire pipeline -- ingestion, embedding, storage, and retrieval -- can run on your own hardware. ### Workflow Engine The workflow engine is a DAG-based (directed acyclic graph) orchestration system with over 30 node types. Workflows are visual, version-controlled, and shareable. Each node in the graph performs a specific operation -- LLM call, conditional branch, data transformation, API request, human approval gate -- and data flows between nodes with type-safe connections. Workflows can combine cloud and local AI providers in a single execution graph. A workflow might use GPT-4o for a summarization step and a local Llama 3 instance for a classification step, routing each job to the appropriate infrastructure automatically. ## Cost Efficiency Running AI workloads on your own hardware dramatically reduces marginal costs once the initial investment is amortized. The following table compares representative cloud API costs against amortized self-hosted costs: | Scenario | Cloud API Cost | Own Hardware Cost | Savings | | ------------------------- | -------------- | ---------------------- | ------- | | 1M tokens/day (GPT-4o) | ~$15/day | ~$0.50/day (amortized) | 97% | | 100 concurrent inferences | ~$3,000/month | ~$500/month | 83% | | 10TB vector storage | ~$230/month | ~$20/month | 91% | These figures assume amortized hardware costs over a typical three-year lifecycle and include electricity and maintenance estimates. Actual savings depend on utilization rates, hardware selection, and workload mix. AceTeam charges a platform fee of 10-15% of compute value for orchestration, management, and platform services. This compares favorably to the 70-90% gross margins that hyperscale cloud providers charge on top of raw compute costs. You pay for the software that makes your hardware useful, not for renting someone else's hardware at a markup. For organizations that already own GPU hardware, the math is straightforward: that hardware is a sunk cost whether or not it is utilized. Running AI workloads on it through AceTeam turns idle capital into productive infrastructure. ## Use Cases ### Enterprise AI Factory Large organizations can pool GPU resources from across departments and locations into a unified AI platform. Engineering, data science, and product teams share a common infrastructure layer with access controls, usage tracking, and cost allocation. The platform handles scheduling, routing, and monitoring so that individual teams do not need to manage infrastructure. ### GPU Marketplace Hardware owners can register GPU nodes with the platform and offer compute capacity to other organizations. The platform handles job routing, billing, and access control. Node owners set pricing and availability; the marketplace matches supply with demand. This creates a decentralized alternative to centralized cloud GPU providers. ### Consultant Multi-Tenant Management Consulting firms and agencies manage AI agents and workflows for multiple clients. Each client exists in an isolated organization with its own mesh network, data, and access controls. Consultants can build agents and workflows in their own environment and deploy them to client environments without co-mingling data or credentials. ### Agentic Workflow Automation Teams automate complex business processes using the workflow engine. Multi-step workflows combine AI inference, data retrieval, conditional logic, human approval gates, and external API calls into repeatable, auditable pipelines. Workflows run on the organization's own hardware, keeping sensitive business data within the network boundary. ### Always-On AI Services Organizations deploy AI agents as persistent services -- customer support bots, internal knowledge assistants, document processing pipelines -- that run continuously on dedicated hardware. Unlike cloud API-based deployments, these services have predictable costs and guaranteed capacity because they run on hardware the organization controls. ## Development Roadmap The platform is currently in production with core functionality operational: mesh networking via Nexus, node management via Citadel, queue-based job distribution, the agent builder, the workflow engine, and multi-provider AI support. Development continues across four phases: **Phase 1 -- Routing Intelligence.** Advanced job routing based on real-time node telemetry, cost optimization, and latency-aware placement. Automatic failover between nodes and providers. **Phase 2 -- Industrial Scale.** Support for large-scale deployments with hundreds of nodes, hierarchical queue topologies, cross-region routing, and enterprise administration features. **Phase 3 -- Compute Economy.** GPU marketplace with billing, pricing, SLA management, and settlement. Organizations can monetize idle hardware by offering it to the network. **Phase 4 -- Global Mesh.** Federation across organizational boundaries, enabling cross-organization compute sharing with cryptographic access controls and auditable job provenance. ## Open Source Key components of the platform are open source: - **Workflow Engine** -- The DAG-based workflow execution engine is available at [github.com/adanomad/workflow-engine](https://github.com/adanomad/workflow-engine). It can be used independently of the AceTeam platform for building and running typed, versioned workflow graphs. - **Citadel CLI** -- The node agent is available at [github.com/aceteam-ai/citadel-cli](https://github.com/aceteam-ai/citadel-cli). Organizations can inspect, audit, and build the binary that runs on their hardware. - **Ace CLI** -- A TypeScript CLI for running AI workflows locally, available at [github.com/aceteam-ai/ace](https://github.com/aceteam-ai/ace). Install via npm and execute workflows from your terminal with multi-provider LLM support. - **AceTeam Nodes** -- The Python workflow node library powering local execution, available at [github.com/aceteam-ai/aceteam-nodes](https://github.com/aceteam-ai/aceteam-nodes). Includes 15+ node types and litellm integration for 100+ LLM providers. Open-sourcing these components reflects a core belief: organizations should be able to verify, modify, and trust the software that runs on their infrastructure. The platform layer provides value through orchestration, management, and marketplace services -- not through opacity. ## Summary AceTeam bridges the gap between raw hardware and managed AI services. The platform separates orchestration from compute, distributes workloads through durable queues, and connects infrastructure through an encrypted mesh network. Organizations keep control of their data and hardware while getting the scheduling, routing, monitoring, and application tooling they need to run AI at scale. The result is an AI platform where the organization -- not a cloud vendor -- decides where data lives, which models run, and how much infrastructure costs. --- ## Fabric Overview Source: https://aceteam.ai/docs/fabric-overview Sovereign AI on hardware you own: run open-weight models and self-hosted AI agents locally, with governance and an audit trail built in. AceTeam Fabric turns machines you already own, GPUs, mini PCs, big-memory Macs, bootable USB nodes, into a private AI mesh. It runs open-weight models and self-hosted AI agents locally, orchestrates a team of agents across whatever mixed hardware you have, and records an audit trail for what ran, where, and on whose data. By default, AceTeam hosts the control plane (the web app, scheduling, orchestration); your compute and your data stay on your own hardware. A fully self-hosted control plane is available on the self-host tenant, for teams that need everything, including the control plane, on their own network. ## Why It Matters Organizations adopt the Fabric for five main reasons: - **Governance and an audit trail:** every job that runs on your Fabric can carry an [AEP](/docs/aep-overview) receipt, recording which agent ran it, on which node, and with what data, for later audit. This is the layer that "just run a local model" does not give you, and it is usually what closes the enterprise deal. - **Heterogeneous orchestration:** route a team of agents across mixed hardware instead of buying one big box. A big-memory Mac serves one model over Ollama, an NVIDIA node serves another over the CUDA engines, and Fabric's capability-tag routing sends each job to the node that can actually run it. - **Data sovereignty:** sensitive data never leaves your network. AI inference, RAG indexing, and workflow execution all happen on machines you own and control. - **Cost control:** use hardware you have already purchased. Avoid per-token cloud pricing by running open-weight models on your own GPUs. - **Compliance:** meet regulatory requirements (HIPAA, GDPR, internal policy) that prohibit sending data to third-party inference providers. ## Architecture The Fabric has three core components: ### Nexus -- Coordination Server Nexus is the VPN coordination server that ties your network together. Built on Headscale (an open-source implementation of the Tailscale control plane), Nexus manages an encrypted mesh network connecting all of your nodes. Every device in your Fabric can reach every other device without opening firewall ports or configuring NAT rules. Nexus runs at `nexus.aceteam.ai` and handles key exchange, node registration, and network policy. ### Citadel -- CLI Agent Citadel is a lightweight CLI agent that runs on each of your compute nodes. It registers the machine with Nexus, joins the mesh VPN, advertises the node's capabilities (GPU type, available models, memory), and listens for incoming jobs. Citadel is the only software you install on your hardware. ### Redis Streams -- Job Queue AceTeam uses Redis Streams as the job queue between the platform and your nodes. When a user runs an agent or triggers a workflow, the platform publishes a job to the appropriate stream. Citadel workers on your nodes consume jobs from the stream, execute them locally, and publish results back. ## Security Model All traffic between nodes travels over the Headscale mesh VPN, encrypted end-to-end with WireGuard. The platform never receives raw data from your nodes -- it only sees job metadata and results you explicitly return. Organization-scoped isolation ensures that each organization's devices, keys, and resources are invisible to other organizations. Nodes registered by your team belong to your organization's Headscale user and cannot be accessed by anyone outside your org. ## Governance and Receipts Running a model locally solves data residency. It does not by itself answer who did what, on which node, with which data, and when, the questions an auditor or a compliance buyer actually asks. That is where the [Agentic Execution Protocol (AEP)](/docs/aep-overview) comes in: every job dispatched through the Fabric's Redis Streams job queue can carry an execution record documenting cost, provenance (which model and which node produced the result), and a governance audit trail. Governance and receipts are the enterprise wedge on top of free local tokens: any runtime can run open-weight models locally, but proving what happened afterward is a separate, harder problem. ## Resource Routing The Fabric distinguishes between two types of resources: - **Fungible resources** are interchangeable. LLM inference is a good example: any node with the right model loaded can handle the request. Fungible jobs go through the Redis Streams queue, tagged by model and capability. The first available node picks up the work. - **Specific resources** are tied to a particular machine. A PostgreSQL database or a file store lives on one node, and requests for that resource must go directly to that node. Specific resources are accessed via direct VPN connections routed through the mesh network. This distinction is what makes heterogeneous orchestration possible: the Fabric does not need one uniform box, it needs the right capability tag on the right node. A big-memory Mac can run one model over Ollama (tagged `llm:mistral`), an NVIDIA node can run a different model over the CUDA engines (tagged `gpu:rtx3090`), and both sit on the same job queue. Add a machine with a new capability and the fleet can serve more, without reconfiguring anything else. ## Next Steps - [Install Citadel](/docs/citadel-setup) on your first node - [Connect additional nodes](/docs/connecting-nodes) and configure resource routing --- ## Citadel Setup Source: https://aceteam.ai/docs/citadel-setup Install the Citadel CLI, authenticate, and register your first compute node. Citadel is the CLI agent that connects your hardware to the AceTeam Fabric. This guide walks you through installing Citadel, authenticating your machine, and verifying that it appears in your Fabric dashboard. ## Prerequisites - An AceTeam account with an active organization - A machine running Linux, macOS, or Windows with administrator access - Outbound internet access (HTTPS and WireGuard UDP) ## Installation ### Linux and macOS Open a terminal and run the install script: ```bash curl -fsSL https://get.aceteam.ai/citadel.sh | sudo bash ``` The script detects your operating system and architecture, downloads the appropriate binary, and places it in `/usr/local/bin/citadel`. ### Windows Open PowerShell as Administrator and run: ```powershell irm https://get.aceteam.ai/citadel.ps1 | iex ``` The installer adds `citadel.exe` to your system PATH. You may need to restart your terminal for the PATH change to take effect. ## Device Authentication Once installed, initialize Citadel to register your machine with the Fabric: ```bash sudo citadel init ``` This starts the device authorization flow: 1. **CLI displays a user code.** You will see a short code on screen, for example `ABCD-1234`, along with a URL. 2. **Open the device auth page.** In your browser, navigate to [aceteam.ai/device](https://aceteam.ai/device). 3. **Log in to AceTeam.** If you are not already signed in, the page will prompt you to log in with your AceTeam credentials. 4. **Enter the user code.** Type or paste the code shown in your terminal. Confirm that you want to authorize the device. 5. **CLI completes registration.** Back in your terminal, Citadel detects the approval, receives its authentication credentials, and registers with the Nexus VPN coordination server. The machine joins your organization's encrypted mesh network. 6. **Node appears in your dashboard.** Open the Fabric page in AceTeam. Your new node should appear in the device list within a few seconds. The entire flow typically takes under a minute. The device code expires after 15 minutes if not used. ## Verify the Connection After initialization completes, check node status: ```bash citadel status ``` You should see output confirming that the node is connected to Nexus, showing its mesh IP address, organization, and connection state. If the status shows "disconnected," verify that your firewall allows outbound UDP traffic on port 41641 (WireGuard). ## Network-Only Mode If you want to join the mesh VPN without running the full Citadel agent (for example, to access other Fabric nodes from a development laptop), use the `--network-only` flag: ```bash sudo citadel init --network-only ``` In network-only mode, the machine joins the encrypted network and can communicate with other nodes, but it does not advertise compute capabilities or accept jobs from the queue. This is useful for administrative machines, jump hosts, or developer workstations that need VPN access without serving as compute nodes. ## Troubleshooting - **"Permission denied"** -- Citadel requires root or administrator privileges for VPN configuration. Use `sudo` on Linux/macOS or run PowerShell as Administrator on Windows. - **Code expired** -- Run `sudo citadel init` again to generate a fresh device code. - **Node not appearing in dashboard** -- Check that you are viewing the correct organization. Nodes are scoped to the organization of the user who authorized them. ## Next Steps - [Connect additional nodes](/docs/connecting-nodes) using preauth keys for unattended setup - Learn about [resource routing](/docs/connecting-nodes#resource-routing) to control how jobs reach your hardware --- ## Coding Agent & File Browser Source: https://aceteam.ai/docs/coding-agent Let your AI agent read, edit, and search code on your own machine. Browse files from the web UI. AceTeam agents can operate directly on code and files stored on your own hardware. The agent reads, edits, and searches files through Citadel -- the same CLI that connects your machine to the Fabric. You can also browse your node's filesystem from the web UI without giving up control of your data. ## Setting Up Citadel for Code Editing If you have already run `citadel init`, your node is authenticated. The next step is to tell Citadel which directory the agent is allowed to work in. ### Start the Worker with a Workspace ```bash citadel work --workspace ~/my-project ``` The `--workspace` flag sets the root directory for all file operations. The agent can read, write, edit, and search files inside this directory. It cannot access anything outside it. You can also set the workspace via environment variable: ```bash export CITADEL_WORKSPACE=~/my-project citadel work ``` If neither the flag nor the variable is set, Citadel defaults to `~/citadel-node/workspace`. ### What the Agent Can Do Once the workspace is configured, your agent has access to these file operations: - **Read** -- Read the contents of any file in the workspace. - **Write** -- Create new files or overwrite existing ones. - **Edit** -- Make targeted edits to specific sections of a file without rewriting the whole thing. - **List** -- Browse directory contents, including nested folders. - **Search** -- Search file contents by pattern across the workspace. These operations run locally on your machine. Files never leave your hardware -- the agent sends instructions over the encrypted VPN, and Citadel executes them in place. ### Example Conversation > You: "Read the main config file in src/config.ts and add a new feature flag called enableBetaUI, defaulting to false." The agent reads the file, finds the right location, and makes the edit. You see the diff in the chat. No cloud IDE, no copy-pasting code into a chat window. ## File Browser The file browser lets you see what is on your compute nodes directly from the AceTeam web app. ### Navigating to Your Node's Files 1. Open **Data** from the main navigation. 2. In the left sidebar, scroll down to the **Compute Nodes** section. 3. Click on your node's name. 4. Browse the workspace directory tree. Click folders to navigate, click files to view contents. The file browser shows the contents of the workspace directory you configured when starting `citadel work`. If you did not set a workspace, it shows the default `~/citadel-node/workspace`. ### What You Can Do - View directory listings with file names, sizes, and modification dates. - Open text files to read their contents. - Navigate through nested folder structures. The file browser is read-only from the web UI. To make changes, use the agent or work on your machine directly. ## Memory Auto-Extraction Your agent learns from your conversations automatically. After each exchange, the agent extracts useful facts, preferences, and patterns from what you discussed and stores them for future reference. This means: - **Facts persist across sessions.** If you tell the agent your project uses TypeScript with strict mode, it remembers that next time. - **Preferences are learned.** If you consistently ask for a specific coding style or naming convention, the agent picks up on the pattern. - **No manual setup.** You do not need to configure anything. Memory extraction runs in the background after each conversation turn. The extracted memories are scoped to your agent and your user account. Other users in your organization do not see your personal memories, and your agent's memories do not leak to other agents. ## Context Compaction Long conversations no longer hit a wall. When a conversation approaches the model's context limit, the agent automatically summarizes older turns into a compact summary and continues the conversation with full awareness of what came before. This means you can have extended working sessions -- debugging a complex issue, iterating on a design, or working through a multi-file refactor -- without the agent losing track of earlier context or the conversation stopping with an error. Compaction happens transparently. You will not see a notification or interruption. The conversation simply keeps going. --- ## Connecting Nodes Source: https://aceteam.ai/docs/connecting-nodes Add compute nodes, manage preauth keys, and configure resource routing. Once your first node is running, you can expand your Fabric by adding more machines. This guide covers preauth keys for unattended registration, node capability tagging, resource routing, monitoring, and node removal. ## Adding Nodes with Preauth Keys The interactive device auth flow works well for a single machine, but it is impractical when provisioning a fleet. Preauth keys let you register nodes without manual browser interaction. ### Generate a Preauth Key From the AceTeam web UI, open the Fabric page, navigate to the Preauth Keys section, and click **Create Key**. Choose an expiration (1 hour to 90 days) and whether the key is single-use or reusable. Copy the key. You can also generate keys from the CLI: ```bash citadel keys create --expiry 24h --reusable ``` ### Register a Node On the target machine, install Citadel and initialize with the preauth key: ```bash sudo citadel init --auth-key ``` The node joins the mesh VPN and registers with the platform immediately, with no browser step required. This makes it straightforward to script Citadel setup in cloud-init, Ansible playbooks, or container entrypoints. ### Key Management - **Single-use keys** are consumed after one registration. Use these for one-off server additions. - **Reusable keys** can register multiple nodes until they expire. Use these for batch provisioning or auto-scaling groups. - Revoke a key at any time from the web UI or with `citadel keys revoke `. Revoking a key does not disconnect nodes already registered with it. ## Node Capabilities Tag each node with the resources it provides so the platform can route work correctly. ```bash citadel config set capabilities llm:llama3,gpu:a100,ram:128gb ``` Capabilities are free-form key-value labels. Common conventions: | Tag | Meaning | | ------------- | --------------------------------- | | `llm:llama3` | Node can serve Llama 3 inference | | `llm:mistral` | Node can serve Mistral inference | | `gpu:a100` | Node has NVIDIA A100 GPU(s) | | `gpu:h100` | Node has NVIDIA H100 GPU(s) | | `ram:256gb` | Node has 256 GB system memory | | `storage:ssd` | Node has SSD-backed local storage | | `db:postgres` | Node runs a PostgreSQL instance | The platform reads these tags when deciding where to route jobs. ## Resource Routing The Fabric uses two routing strategies depending on the resource type: ### Fungible Resources (Queue-Based) Fungible resources are interchangeable -- any node with the matching capability can handle the job. LLM inference is the primary example. When a user sends a message to an agent backed by a self-hosted model, the platform publishes a job to a Redis Streams queue tagged with the required capability (for example, `llm:llama3`). The first available node with that capability picks up the job, runs inference, and returns the result. This approach automatically load-balances across your fleet. If you have three nodes tagged `llm:llama3`, jobs distribute across all three. ### Specific Resources (Direct Connection) Specific resources live on a particular node and cannot be moved. Databases, file stores, and specialized services fall into this category. When the platform needs to access a specific resource, it routes the request directly to the target node over the mesh VPN using its stable mesh IP address. You configure specific resource endpoints in the Fabric dashboard by selecting a node and registering the service (protocol, port, and a human-readable name). ## Monitoring Nodes The Fabric dashboard displays each node's status in real time: - **Online** -- Connected to the mesh, accepting jobs - **Idle** -- Connected but no active jobs - **Busy** -- Currently processing one or more jobs - **Offline** -- Not connected to the mesh Click any node to see detailed metrics: uptime, job history, current capability tags, mesh IP address, and last heartbeat time. From the CLI, run `citadel status --all` to see a summary of every node in your organization. ## Removing Nodes To disconnect a node from the Fabric: ```bash sudo citadel leave ``` This deregisters the machine from Nexus, tears down the VPN tunnel, and removes the node from your dashboard. The Citadel binary remains installed. To fully uninstall, remove the binary and its configuration directory: ```bash sudo citadel leave sudo rm -rf /usr/local/bin/citadel /etc/citadel ``` You can also remove a node from the web UI by selecting it on the Fabric page and clicking **Remove Node**. This forces deregistration even if the node itself is unreachable. ## Next Steps - Review the [Fabric Overview](/docs/fabric-overview) for architecture details - Build an [agent](/docs/first-agent) that runs on your self-hosted infrastructure --- ## GPU Compute Source: https://aceteam.ai/docs/gpu-compute Deploy models, monitor GPUs, benchmark performance, and manage power on your own hardware. The Fabric provides a full set of tools for managing GPU workloads on your own nodes. You can browse a model catalog, deploy models with a single API call, monitor GPU metrics in real time, benchmark inference speed, and configure power management -- all through the dashboard or API. ## Model Marketplace The model marketplace shows a catalog of popular open-weight models with GPU compatibility information for your nodes. For each model and quantization level, the marketplace reports which of your nodes have enough VRAM to run it. ### Browsing Models Open the Fabric page in AceTeam and navigate to the **Models** tab. You will see a list of models organized by family (Llama, Mistral, Gemma, Phi, Qwen, and others). Each entry shows: - **Parameter count** and size class (small, medium, large, xlarge) - **Quantization options** with VRAM requirements (e.g., Q4_K_M at 5 GB, FP16 at 16 GB) - **Compatible nodes** -- which of your online nodes have enough VRAM for each quantization Models are enriched with live data from your Fabric. If no nodes are online, the marketplace still shows the catalog but cannot report compatibility. ### Deploying a Model You can deploy a model from the marketplace or via the API. **From the dashboard:** Click **Deploy** on any compatible model/node combination. The platform starts the inference engine on the selected node. **Via API:** ```bash curl -X POST https://aceteam.ai/api/fabric/provision \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{ "model": "meta-llama/Llama-3-8B-Instruct", "gpu_preference": "any" }' ``` The provisioning endpoint handles the full workflow: 1. Estimates VRAM requirements from the model name 2. Scans your online nodes for one with enough GPU memory 3. Starts the inference engine on the selected node 4. Returns the node ID and status You can also specify a template (`gpu-inference`, `multi-model`, or `cpu-inference`) or let the platform auto-detect the best one from the model name. ### Provisioning Templates Templates define the inference engine, default port, and resource requirements for different deployment scenarios: | Template | Engine | Use Case | | --------------- | --------- | ------------------------------------- | | `gpu-inference` | vLLM | Single large model on a dedicated GPU | | `multi-model` | Ollama | Multiple smaller models sharing a GPU | | `cpu-inference` | llama.cpp | GGUF models on CPU (no GPU required) | List available templates: ```bash curl https://aceteam.ai/api/fabric/provision/templates \ -H "Authorization: Bearer act_your_api_key" ``` ## GPU Dashboard The GPU dashboard provides real-time and historical metrics for every GPU node in your Fabric. ### Real-Time Metrics Each node's detail page shows current GPU status: - **GPU utilization** -- percentage of compute in use - **VRAM usage** -- memory allocated vs. total, per GPU - **Temperature** -- GPU temperature in Celsius ### Historical Graphs View metrics over time with configurable periods: | Period | Resolution | | -------- | --------------------------------- | | 1 hour | Every data point (~30s intervals) | | 6 hours | 1-minute averages | | 24 hours | 5-minute averages | | 7 days | 30-minute averages | **Via API:** ```bash curl "https://aceteam.ai/api/fabric/nodes/{nodeId}/metrics?period=24h" \ -H "Authorization: Bearer act_your_api_key" ``` Returns a time series of `gpu_util_pct`, `vram_used_gb`, `vram_total_gb`, and `temperature_c` data points, downsampled to the appropriate resolution for the requested period. ## Multi-Model Serving Run multiple models on a single GPU using Ollama or a compatible multi-model engine. The API provides endpoints to deploy and remove models and to view VRAM usage per GPU. **Note:** The model list depends on telemetry data from Citadel heartbeats. If your Citadel version does not yet report per-model details, the models list may be empty even when models are running. The VRAM usage data is available when GPU metrics are reported. ### Viewing Running Models ```bash curl https://aceteam.ai/api/fabric/nodes/{nodeId}/models \ -H "Authorization: Bearer act_your_api_key" ``` Returns: - A list of loaded models with their engine, status, and port - GPU information including total and used VRAM per GPU - Overall VRAM budget (total and used across all GPUs) ### Deploying and Removing Models Deploy a model to a node: ```bash curl -X POST https://aceteam.ai/api/fabric/nodes/{nodeId}/models/deploy \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"model_name": "llama3:8b", "engine": "ollama"}' ``` Remove a model: ```bash curl -X DELETE "https://aceteam.ai/api/fabric/nodes/{nodeId}/models?model_name=ollama" \ -H "Authorization: Bearer act_your_api_key" ``` Removing a model stops the engine on the node. Starting a new deployment with a different model reuses the same GPU resources. ## Inference Benchmarks Measure inference speed for a model. Benchmarks report tokens per second, time to first token, and total latency across multiple iterations. **Note:** Benchmark jobs are dispatched to the general GPU queue and run on whichever node picks up the job. The results are associated with the node ID you specify, but execution may occur on a different node if the target node is busy. Results are cached in memory for 1 hour. ### Running a Benchmark ```bash curl -X POST https://aceteam.ai/api/fabric/nodes/{nodeId}/benchmark \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"model_name": "llama3:8b", "num_iterations": 3}' ``` The response includes per-iteration results and an aggregate with median, p95, min, and max for each metric: ```json { "node_id": "my-node", "model_name": "llama3:8b", "aggregate": { "tokens_per_sec": { "median": 42.5, "p95": 38.1, "min": 37.2, "max": 45.3 }, "ttft_ms": { "median": 120, "p95": 145, "min": 110, "max": 150 }, "total_latency_ms": { "median": 2340, "p95": 2580, "min": 2200, "max": 2600 } } } ``` Results are cached for 1 hour. Use `"force": true` to re-run. Each organization is limited to 5 benchmark runs per hour. ### Benchmark Leaderboard Compare models across your nodes: ```bash curl "https://aceteam.ai/api/fabric/benchmarks/leaderboard?model_name=llama3:8b" \ -H "Authorization: Bearer act_your_api_key" ``` Returns all benchmark results for the specified model, sorted by tokens per second (fastest first). Omit `model_name` to see all results. ## Scale-to-Zero Configure GPU services to auto-suspend after a period of inactivity and resume on the next request. This saves power and reduces wear on GPUs that are not continuously serving traffic. ### Configuring Auto-Suspend Set the idle timeout for a service on a node: ```bash curl -X PUT https://aceteam.ai/api/fabric/nodes/{nodeId}/services/ollama/idle-config \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"auto_suspend_minutes": 15}' ``` Set `auto_suspend_minutes` to `0` to disable auto-suspend (always-on mode). Valid values range from 5 to 1440 minutes (24 hours). ### Checking Status ```bash curl https://aceteam.ai/api/fabric/nodes/{nodeId}/services/ollama/idle-config \ -H "Authorization: Bearer act_your_api_key" ``` Returns the current configuration, service state (`active` or `suspended`), last request timestamp, and any cached models. ### Manual Resume If a service is suspended, you can wake it immediately: ```bash curl -X POST https://aceteam.ai/api/fabric/nodes/{nodeId}/services/ollama/resume \ -H "Authorization: Bearer act_your_api_key" ``` Resume dispatches a service start command to the node. If the cold start takes longer than expected, the endpoint returns a 503 with a `Retry-After` header. Incoming inference requests to a suspended service automatically trigger a resume. The first request after suspension may experience higher latency while the model reloads. ## Node Reliability Scoring Every node in your Fabric receives a reliability score based on its operational history. The score helps you identify which nodes are most dependable for production workloads. ### Score Components The reliability score (0-100) is a weighted combination of four factors: | Factor | What it measures | | ----------------------- | ------------------------------------------------------------ | | **Uptime** | Percentage of time the node was online | | **Completion rate** | Fraction of jobs that completed successfully | | **Latency** | Average job processing time | | **Heartbeat stability** | Consistency of status reports (gaps indicate network issues) | Nodes are categorized into tiers: | Tier | Score | Meaning | | ------ | -------- | ------------------------------------- | | Green | 95+ | Highly reliable | | Yellow | 80-94 | Generally reliable, occasional issues | | Red | Below 80 | Needs attention | ### Viewing Reliability ```bash # Current score for a node (default: last 30 days) curl "https://aceteam.ai/api/fabric/nodes/{nodeId}/reliability?period_days=30" \ -H "Authorization: Bearer act_your_api_key" # Daily history curl "https://aceteam.ai/api/fabric/nodes/{nodeId}/reliability/history?days=30" \ -H "Authorization: Bearer act_your_api_key" # Rank all nodes by reliability curl "https://aceteam.ai/api/fabric/reliability/ranking" \ -H "Authorization: Bearer act_your_api_key" ``` The ranking endpoint returns all nodes in your organization sorted by score, making it easy to identify which nodes to prioritize for critical workloads. ## Model Caching Register model weights for pre-pulling to your nodes. Once cached, models load from local storage instead of downloading on first request. **Note:** The cache registry tracks which models should be present on each node. The actual download is handled by the Citadel node's inference engine. If a model is registered but not yet downloaded by the engine, the first inference request may still trigger a download. ### Caching a Model ```bash curl -X POST https://aceteam.ai/api/fabric/nodes/{nodeId}/models/cache \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"model_name": "llama3:8b", "engine": "ollama"}' ``` The cache pull is dispatched asynchronously -- the endpoint returns a job ID immediately. Large models (e.g., 70B parameter models) may take several minutes to download depending on network speed. ### Listing Cached Models ```bash curl https://aceteam.ai/api/fabric/nodes/{nodeId}/models/cache \ -H "Authorization: Bearer act_your_api_key" ``` ### Removing a Cached Model ```bash curl -X DELETE "https://aceteam.ai/api/fabric/nodes/{nodeId}/models/cache?model_name=llama3:8b" \ -H "Authorization: Bearer act_your_api_key" ``` ## GPU Power Management Hibernate and wake GPU services on your nodes. Hibernation stops inference engines to save power while keeping the node on the VPN mesh and responsive to management commands. ### Hibernate ```bash curl -X POST https://aceteam.ai/api/fabric/nodes/{nodeId}/power/hibernate \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"services": ["vllm", "ollama"]}' ``` If no `services` list is provided, all inference engines (vllm, ollama, llamacpp) are stopped. ### Wake ```bash curl -X POST https://aceteam.ai/api/fabric/nodes/{nodeId}/power/wake \ -H "Authorization: Bearer act_your_api_key" ``` ### Power State Check the current power state: ```bash curl https://aceteam.ai/api/fabric/nodes/{nodeId}/power \ -H "Authorization: Bearer act_your_api_key" ``` Returns: ```json { "node_id": "my-node", "power_state": "active", "config": { "auto_hibernate_minutes": 30, "enabled": false } } ``` ### Auto-Hibernate Configuration Save your preferred auto-hibernate settings for a node. When enabled, the node will automatically hibernate after the specified idle period. **Note:** Auto-hibernate configuration is saved but enforcement is not yet active. Use the manual hibernate/wake controls above for immediate power management. Automatic enforcement based on idle detection is coming soon. ```bash curl -X PUT https://aceteam.ai/api/fabric/nodes/{nodeId}/power/config \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"auto_hibernate_minutes": 30, "enabled": true}' ``` ## Next Steps - [Install Citadel](/docs/citadel-setup) on your first GPU node - [Connect additional nodes](/docs/connecting-nodes) and configure capabilities - [Browse the model catalog](/docs/gpu-compute#model-marketplace) and deploy your first model - Set up [ACET tokens](/docs/acet-tokens) for compute billing --- ## Sandboxes Source: https://aceteam.ai/docs/sandboxes Run ephemeral Docker containers on your own hardware. Execute code, checkpoint state, and build custom environments. Sandboxes are ephemeral Docker containers that run on your Citadel nodes. They give agents and users an isolated environment to execute code, install packages, and run experiments -- all on hardware you control. Data never leaves your network. ## How It Works When you create a sandbox, AceTeam dispatches a Docker command to one of your Citadel nodes. The container starts with network isolation (`--network none`) and stays alive for interactive use. You can run commands inside it, checkpoint its state to a Docker image, and destroy it when done. ``` AceTeam Platform Your Citadel Node | | | "Create sandbox" | |------------------------------>| | | docker run -d --network none ubuntu:24.04 | | | "Run: ls /workspace" | |------------------------------>| | | docker exec sandbox-abc12345 sh -c "ls /workspace" | | | "Checkpoint" | |------------------------------>| | | docker commit sandbox-abc12345 sandbox-abc12345:checkpoint-... ``` ## Prerequisites - A Citadel node with Docker installed (included by default when you run `citadel init`) - The node must be online and connected to the Fabric ## Using Sandboxes via MCP The primary way to interact with sandboxes is through MCP tools -- available in Claude Code, Cursor, and other MCP-compatible AI coding tools. ### Creating a Sandbox ``` sandbox_create(node_id="my-node", image="python:3.12-slim") ``` Returns the sandbox ID, status, and image. The default image is `ubuntu:24.04`. ### Running Commands ``` sandbox_exec(node_id="my-node", sandbox_id="sandbox-abc12345", command="python --version") ``` Returns stdout, stderr, and exit code. The optional `workdir` parameter sets the working directory inside the container. **Note:** Sandboxes run with `--network none` (no network access). Commands that require internet (like `pip install` or `apt-get update`) will fail inside a running sandbox. To pre-install dependencies, use a [custom template](#custom-templates) that bakes them into the image at build time, when network access is available. ### Checkpointing Save the current state of a sandbox as a Docker image. This preserves installed packages, created files, and any other container changes. ``` sandbox_checkpoint(node_id="my-node", sandbox_id="sandbox-abc12345") ``` Returns the image tag (e.g., `sandbox-abc12345:checkpoint-20260613143022`). The checkpoint is stored locally on the Citadel node. Create a new sandbox from a checkpoint by passing its image tag: ``` sandbox_create(node_id="my-node", image="sandbox-abc12345:checkpoint-20260613143022") ``` ### Listing Sandboxes ``` sandbox_list(node_id="my-node") ``` Returns a table of all sandbox containers on the node, including their status and image. ### Destroying a Sandbox ``` sandbox_destroy(node_id="my-node", sandbox_id="sandbox-abc12345") ``` This is irreversible. Use checkpoint first if you want to preserve the sandbox state. ## Custom Templates Custom templates let you define reusable sandbox environments with your tools pre-installed. Instead of starting from a generic image every time, create a template once and spawn sandboxes from it. Templates are built with network access, so they can download packages and dependencies. The resulting sandboxes then run without network access. ### Creating a Template via MCP **Option 1: Base image + setup commands** ``` sandbox_create_template( name="ml-research", description="Python ML environment", base_image="python:3.12-slim", setup_commands=["pip install numpy pandas scikit-learn matplotlib"] ) ``` **Option 2: Raw Dockerfile** ``` sandbox_create_template( name="node-dev", dockerfile_content="FROM node:20-slim\nRUN npm install -g typescript ts-node" ) ``` ### Creating a Template via API Templates also have a REST API: ```bash curl -X POST https://aceteam.ai/api/sandbox/templates \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{ "name": "ml-research", "base_image": "python:3.12-slim", "setup_commands": [ "pip install numpy pandas scikit-learn matplotlib" ] }' ``` The template is built asynchronously on one of your Citadel nodes. The response includes the template status (`building`, `ready`, or `failed`) and an image tag. ### Listing Templates Via MCP: ``` sandbox_list_templates() ``` Via API: ```bash curl https://aceteam.ai/api/sandbox/templates \ -H "Authorization: Bearer act_your_api_key" ``` Returns both built-in templates and your custom templates. Built-in templates are always available and cannot be deleted. ### Using a Template Once a template status is `ready`, create sandboxes from its image: ``` sandbox_create(node_id="my-node", image="aceteam-sandbox:custom-abc12345") ``` The image tag is returned when you create the template. ### Deleting a Template ```bash curl -X DELETE https://aceteam.ai/api/sandbox/templates/{templateId} \ -H "Authorization: Bearer act_your_api_key" ``` Built-in templates cannot be deleted. ## Security Sandbox containers run with `--network none` by default, meaning they have no network access. This prevents code running inside a sandbox from making outbound requests or accessing other services on your network. The sandbox ID format (`sandbox-<8 hex chars>`) is validated on every request to prevent injection attacks. All commands are dispatched through the Citadel job system over the encrypted VPN mesh. ## MCP Tool Reference | Tool | Description | | ------------------------- | ------------------------------------ | | `sandbox_create` | Create a sandbox on a Citadel node | | `sandbox_exec` | Run a command inside a sandbox | | `sandbox_list` | List sandboxes on a node | | `sandbox_checkpoint` | Save sandbox state as a Docker image | | `sandbox_destroy` | Remove a sandbox container | | `sandbox_create_template` | Create a custom template | | `sandbox_list_templates` | List available templates | ## Next Steps - [Install Citadel](/docs/citadel-setup) on the machine where you want to run sandboxes - Use the [Coding Agent](/docs/coding-agent) to work with files on your nodes - Set up [GPU Compute](/docs/gpu-compute) to run AI models alongside your sandboxes - Use the [SDKs](/docs/sdks) for hosted instance management with safety enforcement --- ## ACET Compute Tokens Source: https://aceteam.ai/docs/acet-tokens Purchase ACET tokens to pay for GPU inference on the Sovereign Compute Fabric. Programmatic registration for agents. ACET tokens are the billing currency for GPU inference on the Sovereign Compute Fabric. When you run models on your own hardware or use GPU nodes shared within your organization, ACET tokens track usage and compensate node operators. ## How ACET Works - **1 ACET = $0.001 USD** (fixed peg) - Tokens are purchased with USD and credited to your organization's ledger - When your agents or workflows run inference on the Fabric, ACET is debited from your balance - Revenue is split between the node operator (80%) and the platform (20%) ### Pricing by Model Size | Size Class | Parameter Range | ACET per 1K Tokens | | ---------- | --------------- | ------------------ | | Small | 0-8B | 1 | | Medium | 8B-70B | 5 | | Large | 70B-400B | 25 | | XLarge | 400B+ | 100 | ## Purchasing ACET Tokens ### Via the Dashboard Go to **Settings > Billing** and click **Buy ACET Tokens**. Select an amount (minimum 1,000 ACET / $1.00) and complete checkout through Stripe. After successful payment, tokens are credited to your organization's ledger immediately. ### Via API Create a Stripe Checkout session programmatically: ```bash curl -X POST https://aceteam.ai/api/acet/checkout \ -H "Authorization: Bearer act_your_api_key" \ -H "Content-Type: application/json" \ -d '{"amount_acet": 10000}' ``` Response: ```json { "checkout_url": "https://checkout.stripe.com/c/pay/...", "session_id": "cs_live_...", "amount_acet": 10000, "amount_usd": 10.0, "current_balance": 5000 } ``` Redirect your user to `checkout_url` to complete the payment. After payment, they are returned to the billing page. The `amount_acet` must be a multiple of 1,000. ### Checking Your Balance ```bash curl https://aceteam.ai/api/acet/balance \ -H "Authorization: Bearer act_your_api_key" ``` ### Transaction History ```bash curl https://aceteam.ai/api/acet/history \ -H "Authorization: Bearer act_your_api_key" ``` Returns a paginated list of all ledger entries -- purchases, inference debits, and operator credits. ## Agent Self-Registration Agents and automated systems can programmatically create an AceTeam organization and receive an API key, without any manual signup flow. ### How It Works The registration endpoint creates: 1. A user account associated with the provided email 2. An organization 3. A read-scoped API key for the `/api/fabric/**` and `/api/whoami` endpoints ```bash curl -X POST https://aceteam.ai/api/fabric/register \ -H "Content-Type: application/json" \ -d '{ "email": "agent@example.com", "org_name": "My Agent Org" }' ``` Response: ```json { "organization_id": "uuid-...", "api_key": "act_xxxxxxxxxxxx", "aceteam_url": "https://aceteam.ai" } ``` The returned API key is scoped to read-only access on Fabric endpoints. To unlock full API access (write, execute, additional endpoints), verify the email address or create a new key from the dashboard. ### Rate Limits Self-registration is rate-limited to prevent abuse: - **Per IP:** Limited registrations per hour - **Per email domain:** Limited registrations per domain per hour (prevents x-forwarded-for spoofing) If you hit a rate limit, the response includes a `Retry-After` header with the number of seconds to wait. ### Use Cases - **CI/CD pipelines** that need to provision compute resources programmatically - **Agent frameworks** that bootstrap their own infrastructure - **Multi-tenant platforms** where each tenant gets their own isolated AceTeam organization ## Revenue Split When ACET tokens are spent on inference, the revenue is split: | Recipient | Share | Description | | ------------- | ----- | ----------------------------------------- | | Node operator | 80% | The organization running the GPU hardware | | Platform | 20% | AceTeam infrastructure and orchestration | The operator always receives `ceil(80%)` of the ACET cost, ensuring small jobs never short-change the hardware provider. Operators can view their earnings on the Fabric dashboard under **Earnings**, or via the API: ```bash # Total earnings across all nodes curl https://aceteam.ai/api/fabric/earnings \ -H "Authorization: Bearer act_your_api_key" # Earnings for a specific node curl https://aceteam.ai/api/fabric/nodes/{nodeId}/earnings \ -H "Authorization: Bearer act_your_api_key" ``` ## Next Steps - [Deploy a model](/docs/gpu-compute#model-marketplace) on your GPU nodes - [Set up Citadel](/docs/citadel-setup) to contribute your hardware to the Fabric - Review [Billing & Credits](/docs/billing) for platform-level billing (separate from ACET) --- ## Citadel OS Source: https://aceteam.ai/docs/citadel-os Pre-built VM image with GPU drivers, Docker, and Citadel pre-installed. Deploy a GPU node in minutes. Citadel OS is a pre-built Ubuntu 24.04 VM image designed for zero-configuration GPU node onboarding. Instead of manually installing NVIDIA drivers, Docker, and the Citadel CLI on each machine, you deploy the Citadel OS image and provide an auth key -- the node joins your Fabric automatically on first boot. ## What's Included The image ships with everything needed to serve GPU inference: | Component | Version | Purpose | | ------------------------ | --------- | --------------------------------------------------------------- | | Ubuntu | 24.04 LTS | Base operating system | | NVIDIA Driver | 570.x | GPU hardware support | | CUDA Toolkit | 12.8 | GPU compute libraries (DKMS-registered) | | Docker CE | Latest | Container runtime | | NVIDIA Container Toolkit | Latest | GPU access inside Docker containers (nvidia as default runtime) | | Citadel CLI | Latest | Fabric agent and VPN client | | vLLM | Latest | Pre-pulled inference engine image | The image is approximately 20 GB. Build time is 20-40 minutes depending on network speed (CUDA and vLLM downloads dominate the build). ## Building the Image Citadel OS images are built with [Packer](https://developer.hashicorp.com/packer/install), producing a QCOW2 disk image suitable for import into Proxmox, QEMU, or other KVM-based hypervisors. ### Prerequisites - Packer >= 1.9 - QEMU/KVM (`apt install qemu-system-x86 qemu-utils`) - KVM access (`/dev/kvm`) - ~20 GB free disk space ### Build ```bash cd citadel-cli/packer/ # Initialize Packer plugins (first time only) packer init citadel-node.pkr.hcl # Build the image packer build citadel-node.pkr.hcl ``` The output image is written to `output/citadel-node.qcow2`. You can customize the build with variables: ```bash packer build \ -var 'disk_size=100G' \ -var 'memory=8192' \ -var 'cpus=8' \ citadel-node.pkr.hcl ``` ## Deploying to Proxmox A deployment script handles the full workflow: upload the image, create a VM template, clone a new VM, inject the auth key, and start it. ```bash cd citadel-cli/packer/ ./deploy-to-proxmox.sh \ --host root@pve.local \ --authkey tskey-auth-xxxxx ``` ### Full Options ```bash ./deploy-to-proxmox.sh \ --host root@pve.local \ --storage local-lvm \ --template-id 9000 \ --vm-id 200 \ --vm-name gpu-worker-01 \ --cores 8 \ --memory 32768 \ --authkey tskey-auth-xxxxx ``` Preview commands without executing: ```bash ./deploy-to-proxmox.sh --dry-run \ --host root@pve.local \ --authkey tskey-auth-xxxxx ``` ### What the Script Does 1. Uploads the QCOW2 image to the Proxmox host via SCP 2. Creates a VM and imports the disk, then converts it to a template 3. Clones the template to a new full VM 4. Injects the auth key via a cloud-init snippet 5. Starts the VM ### GPU Passthrough For GPU workloads, pass a physical GPU through to the VM after cloning: ```bash # Add PCI passthrough (adjust the address for your hardware) qm set 200 --hostpci0 0000:01:00,pcie=1 ``` IOMMU must be enabled in your BIOS and Proxmox boot parameters: ```bash # Intel CPUs: add to /etc/default/grub GRUB_CMDLINE_LINUX_DEFAULT="intel_iommu=on" # AMD CPUs GRUB_CMDLINE_LINUX_DEFAULT="amd_iommu=on" ``` ## First Boot On first boot, a one-shot systemd service (`citadel-firstboot.service`) runs automatically: 1. Reads the auth key from `/etc/citadel/authkey` (written by cloud-init) 2. Runs `citadel init --authkey ` to join the mesh network 3. Enables and starts the Citadel worker service 4. Removes the auth key file and disables itself If no auth key is present, the service logs a warning and exits. You can initialize manually later: ```bash ssh citadel@ citadel init --authkey sudo systemctl enable --now citadel-worker.service ``` After the first boot completes, the node appears in your Fabric dashboard and is ready to accept jobs. ## Provisioning Scripts The image is built by running these scripts in sequence during the Packer build: | Script | What it installs | | ----------------- | ---------------------------------------------------------------------- | | `01-base.sh` | System essentials: curl, git, jq, htop, tmux, build-essential, dkms | | `02-nvidia.sh` | NVIDIA driver + CUDA 12.8 via runfile with DKMS | | `03-docker.sh` | Docker CE + NVIDIA Container Toolkit (nvidia as default runtime) | | `04-citadel.sh` | Latest Citadel CLI, systemd worker service (disabled until first boot) | | `05-vllm.sh` | Pre-pulls the `vllm/vllm-openai:latest` Docker image | | `06-firstboot.sh` | Installs the one-shot first-boot service for auth key initialization | ### Customization **Different CUDA version:** Edit `scripts/02-nvidia.sh` and update `CUDA_VERSION` and `DRIVER_VERSION`. See the [CUDA archive](https://developer.nvidia.com/cuda-toolkit-archive) for available versions. **Skip vLLM pre-pull:** Comment out the `05-vllm.sh` provisioner in `citadel-node.pkr.hcl` for a smaller image. You can pull vLLM later when you deploy a model. **Different base image:** Change the `ubuntu_image_url` variable. The provisioning scripts assume Ubuntu/Debian package management. ## Next Steps - Generate a [preauth key](/docs/connecting-nodes#adding-nodes-with-preauth-keys) for unattended deployment - [Deploy a model](/docs/gpu-compute#deploying-a-model) to your new node - View [GPU metrics](/docs/gpu-compute#gpu-dashboard) on the Fabric dashboard --- ## Share Fabric Nodes, List Them on the Marketplace, and Track Earnings Source: https://aceteam.ai/docs/fabric-marketplace-node-sharing Grant another organization access to a compute node, list a verified node on the AceTeam marketplace, set your own pricing, and track earnings and withdrawals. Once a node is connected to your [Sovereign Compute Fabric](/docs/fabric-overview), you can keep it private to your organization, share it with a specific partner organization, or list it on the marketplace so other organizations can route work to it. These are three separate, layered decisions, and each one grants a different, narrower set of capabilities than the last. ## Sharing a Node with Another Organization `fabric_share_node` grants a specific organization access to one of your nodes, without giving up ownership. Only an owner or admin of the organization that owns the node can grant a share, and only into an organization they are themselves an active member of. Sharing grants: - Inference pinning: the other organization can pin an agent's inference to your node and model. - Ping, speed test, and status reads. - The shared file and code read surface for that node. Sharing does not grant shell access. Running a command on your node still requires a separate `node:exec`-scoped API key and a human-approved execution grant, regardless of any sharing in place. `fabric_unshare_node` revokes the grant; both tools are idempotent, so sharing an already-listed organization, or unsharing one that is not listed, is a clean no-op. A node you only have shared access to, rather than own, cannot be re-shared to a third organization. ## Verifying a Node Before Going Live Before a node can be listed on the marketplace, it has to pass hardware verification. `fabric_verify_node` runs a short benchmark (currently against a small reference model) and records the node's GPU VRAM and measured tokens per second. Results are cached, so re-running verification is only needed after a hardware change; pass `force` to re-verify anyway. `fabric_set_node_models` sets which models a node advertises once it is live, so browsing the marketplace shows only models the node is actually configured to serve. ## Listing on the Marketplace `fabric_toggle_marketplace` lists or delists a node. Listing requires the node to already be verified; an unverified node is refused with a pointer back to `fabric_verify_node`. Marketplace visibility is discovery, not a grant. Listing a node makes it discoverable through `fabric_list_nodes` with `marketplace` set to true (which strips org attribution from the result) and lets it serve the fleet wide capability tag queues it opts into itself. It never lets an outside organization address your node directly the way an explicit share does; only ownership or `fabric_share_node` does that. There is no rental or lease record in the schema, so an explicit share is the entire cross org access model, and marketplace listing sits alongside it as a separate, narrower kind of visibility. ## Pricing `fabric_set_pricing` sets a custom ACET rate for a node you own, in ACET per 1,000 tokens: a `base_rate` for every model on the node, plus optional `model_overrides` for specific models. Leaving pricing unset falls back to the platform default rate. `fabric_get_pricing` reads the current pricing for a node you own or that is shared with you; only the owning organization can change it. ## Anonymized Earnings Ladder `fabric_ladder_opt_in` opts a node into a ranked ladder of anonymized earnings across providers. Nobody but you ever sees your node's id, name, or exact earnings; other providers see only a coarse earnings band and a server hashed handle. You must opt in at least one of your own nodes before you can see other providers' rows at all, and the ladder itself stays withheld until enough distinct providers have joined. Opting out removes your node from the ladder immediately. ## Reading Earnings `fabric_node_earnings` and `fabric_earnings_summary` report ACET earned over a day, week, or month, derived from recorded job duration and token counts. `fabric_node_earnings` returns one node's breakdown; `fabric_earnings_summary` returns every node in your organization with per-node and total rows. ## Requesting a Withdrawal `fabric_request_withdrawal` requests a payout of your earned ACET by bank transfer or crypto destination. Only an organization owner or admin can request a withdrawal. The request is validated against your actual available balance (earned minus already pending minus already withdrawn, so a second request cannot double spend the same earnings) and against a minimum payout threshold before it is created. A valid request is written as a real, persisted `pending` record and held for manual review; it is not merely accepted and discarded. `fabric_list_withdrawals` lists your organization's withdrawal requests with their current status, amount, method, and timestamps, most recent first. --- ## Workflow Editor Source: https://aceteam.ai/docs/workflow-editor Navigate the visual DAG editor to build multi-step automations. The workflow editor is where you design automations as visual graphs. Each workflow is a directed acyclic graph (DAG): a set of nodes connected by edges where data flows in one direction, from left to right, with no loops. This model ensures that every workflow has a clear start, predictable execution order, and a defined end. ## Opening the Editor Navigate to **Workflows** in the sidebar and click **New Workflow**, or open an existing workflow from the list. The editor loads a canvas where you can build and modify your automation. ## Canvas Navigation The canvas is your workspace. Use these controls to move around: - **Pan**: Click and drag on empty space to move the canvas, or hold the middle mouse button. - **Zoom**: Scroll up to zoom in, scroll down to zoom out. You can also use the zoom controls in the bottom-left corner of the canvas. - **Fit to view**: Click the fit button to automatically zoom and center the canvas so all nodes are visible. - **Minimap**: A small overview in the corner shows your current viewport relative to the full workflow. Click anywhere on the minimap to jump to that area. ## Adding Nodes The node browser sits on the left side of the editor. Nodes are organized into color-coded categories: **AI** (purple), **Logic** (blue), **Data** (green), **Documents** (orange), **Communication** (teal), **Sales & Outreach** (amber), and more. Each category has a distinct color so you can visually scan for the type of node you need. To add a node: 1. Open the node browser and browse by category, or type in the search bar to filter by name. 2. Click and drag a node onto the canvas, or click it once to place it at the center of your current view. 3. The node appears on the canvas with default settings, ready to be configured. You can also right-click on empty canvas space to open a quick-add menu with search. ## Connecting Nodes with Edges Nodes have input ports on their left side and output ports on their right side. To connect two nodes: 1. Hover over an output port on the source node. The port highlights when it is ready. 2. Click and drag from the output port to an input port on the target node. 3. Release to create the connection. A line (edge) appears between the two nodes. Data flows along edges from the source node's output to the target node's input. A node will not execute until all of its upstream nodes have completed and passed their data forward. To remove a connection, click on the edge to select it and press **Delete**, or right-click the edge and choose **Remove**. ## The DAG Model Because workflows are directed acyclic graphs, you cannot create circular connections. If you try to draw an edge that would create a loop, the editor will reject it. This constraint guarantees that every workflow terminates and that execution order is always deterministic. Nodes with no upstream dependencies execute first. When a node finishes, all downstream nodes that have received their required inputs begin executing. Multiple branches can run in parallel if they do not depend on each other. ## Editing Node Properties Click on any node to open its configuration panel on the right side of the editor. Each node type has different settings, for example, an AI node lets you pick a model and write a prompt, while an HTTP node lets you set the URL, method, and headers. Changes save automatically as you type. You can also rename a node by double-clicking its title on the canvas. Clear names make complex workflows easier to understand at a glance. ## Saving and Naming Workflows Workflows auto-save as you edit. The current save status is shown in the top bar: look for "Saved" or "Saving..." next to the workflow name. To rename a workflow, click its name in the top bar and type a new one. Give workflows descriptive names that reflect their purpose, such as "Weekly Report Generator" or "Lead Qualification Pipeline." You can also add a description in the workflow settings panel to document what the workflow does and when it should run. ## Sticky Notes You can place sticky notes anywhere on the canvas to annotate your workflow. Right-click on empty space and select **Add Sticky Note**, or use the keyboard shortcut. Sticky notes are freeform text areas: use them to document design decisions, leave instructions for collaborators, or mark sections of a complex workflow. They do not affect execution and are purely visual aids. ## Pre-Run Integration Check Before executing a workflow, open the pre-run integration check panel from the toolbar. The panel scans every node in your graph and verifies that all connected services are properly configured: - **OAuth tokens**: checks that Google, Twilio, and other OAuth connections are still valid. - **API keys**: verifies that configured API keys have not expired or been revoked. - **Webhook endpoints**: confirms that target URLs are reachable. - **Required fields**: flags any nodes with missing required configuration. Fix any issues the panel surfaces before running to avoid mid-execution failures. The check runs automatically when you click **Run** if you have not run it manually. ## Next Steps - Learn about the [node types](/docs/workflow-nodes) available in the editor - Set up [triggers](/docs/workflow-triggers) to start your workflows automatically - Understand how to [run and monitor](/docs/workflow-execution) workflow execution --- ## Workflow Nodes Source: https://aceteam.ai/docs/workflow-nodes Understand the node types available in the workflow editor. Nodes are the building blocks of every workflow. Each node performs a single operation: calling an AI model, making an HTTP request, running code, or branching based on a condition. You chain nodes together to create complex automations from simple, composable pieces. ## Node Anatomy Every node has **inputs** on its left side and **outputs** on its right side. Inputs receive data from upstream nodes. Outputs pass results to downstream nodes. Some nodes have a single input and output, while others have multiple (for example, a Conditional node has separate outputs for each branch). ## Node Types ### Trigger Trigger nodes define how a workflow starts. Every workflow must have exactly one trigger node. Triggers produce the initial data that flows into the rest of the graph. See [Workflow Triggers](/docs/workflow-triggers) for details on each trigger type. ### AI AI nodes send a prompt to a language model and return the response. Configuration options include: - **Model**: Select from available providers (OpenAI, Anthropic Claude, Google Gemini, Deepseek, and others). - **System prompt**: Set the model's role and instructions. - **User prompt**: Define the message sent to the model. Use template variables like `{{input.text}}` to inject data from upstream nodes. - **Temperature**: Control response creativity (0 for deterministic, higher for more varied). - **Max tokens**: Limit the length of the response. The AI node outputs the model's response as a string, which you can pass to any downstream node. ### Conditional Conditional nodes branch your workflow based on rules. Define one or more conditions using the data available from upstream nodes. Each condition gets its own output port, plus a fallback "else" port for cases where no condition matches. Conditions support standard comparisons: equals, not equals, contains, greater than, less than, is empty, and regular expression matching. You can combine multiple conditions with AND/OR logic. ### HTTP HTTP nodes make external API calls. Configure the request with: - **URL**: The endpoint to call. Supports template variables. - **Method**: GET, POST, PUT, PATCH, or DELETE. - **Headers**: Key-value pairs, commonly used for authentication tokens. - **Body**: JSON, form data, or raw text for POST/PUT/PATCH requests. - **Timeout**: Maximum time to wait for a response (default is 30 seconds). The node outputs the response status code, headers, and body. You can parse JSON responses in downstream nodes. ### Code Code nodes run custom JavaScript or Python. Use them when built-in nodes do not cover your logic. The code editor provides: - **Language selection**: Choose JavaScript or Python. - **Input access**: Reference upstream data via the `input` object. - **Return value**: Whatever your code returns becomes the node's output. Code nodes run in a sandboxed environment with a 60-second timeout. They have access to standard library functions but cannot make network requests directly. Use an HTTP node for that. ## Data Types Data flowing between nodes falls into these types: - **String**: Text values, including AI model responses and raw HTTP bodies. - **Number**: Integers and decimals, used in calculations and comparisons. - **Sequence**: Ordered lists of items. Downstream nodes can process sequences one item at a time or as a whole. - **File**: Binary data such as uploaded documents or images, referenced by a storage URL. ## Type Casting When a node receives data that does not match its expected input type, the editor applies automatic type casting where possible. Numbers convert to strings, strings that contain valid numbers convert to numbers, and single items convert to single-element sequences. If automatic conversion is not possible, the node will surface an error at execution time. You can also add explicit type conversion by inserting a Code node between two nodes that need different types. This gives you full control over formatting and parsing. ### SubWorkflow SubWorkflow nodes let you embed one workflow inside another. Select an existing workflow and map inputs from the parent graph into the child workflow's trigger. The child workflow runs to completion and returns its output to the parent. **SubWorkflowForEach** and **SubWorkflowParallelForEach** variants iterate over a sequence, running the child workflow once per item, sequentially or in parallel. ### ParallelForEach ParallelForEach processes every item in a sequence concurrently instead of one at a time. It accepts a sequence input, fans out to run the configured downstream branch for each item simultaneously, and collects the results when all branches complete. Use this for batch operations like enriching a list of leads or generating summaries for multiple documents. ### FormulaCalc FormulaCalc evaluates spreadsheet-style formulas against your workflow data. Write expressions using standard operators, math functions, and references to upstream node outputs. Use it for unit conversions, scoring calculations, weighted averages, and other arithmetic that would otherwise require a Code node. ### LookupTable LookupTable maps input values to output values using a configurable reference table. Define rows of key-value pairs and the node returns the matching value for a given input. Useful for category mapping, status code translation, tier-based pricing lookups, and similar reference data patterns. ### Email & Gmail **EmailNode** sends email via SMTP. Configure the recipient, subject, and body, with template variable support for dynamic content. **GmailSendNode** sends email through a connected Google account using OAuth, with full access to Gmail-specific features like labels and threading. ### Google Calendar Four nodes for Google Calendar integration: - **ListCalendarEvents**: retrieve events for a date range. - **CreateCalendarEvent**: create a new event with title, time, attendees, and description. - **CheckCalendarAvailability**: check free/busy status for a set of calendars. - **DeleteCalendarEvent**: remove an event by ID. ### Document Nodes - **ReadDocx**: extract text and structure from DOCX files. - **DoclingConverter**: convert documents between formats using Docling. - **ExcelReader**: read data from Excel spreadsheets into sequences. ### Sales & Outreach Purpose-built nodes for B2B sales automation: - **LeadEnrich**: pull company and contact data from enrichment APIs. - **ColdEmail**: draft and send personalized outreach emails. - **CRMPush**: sync records to your CRM system. ### Calculation & Formatting - **SpreadsheetCalc**: run multi-cell spreadsheet calculations. - **TableBuilder**: construct tabular data from workflow inputs. - **NumberFormat**: format numbers with locale-aware precision, currency, and percentage styles. - **DateCalc**: perform date arithmetic (add/subtract days, calculate differences, format dates). ### Utility Nodes - **TextJoin**: concatenate multiple text inputs with a configurable separator. - **Placeholder**: a no-op node for reserving a spot in your graph during design. - **ExpandJSON**: flatten nested JSON objects into individual output fields. - **SortJSON**: sort JSON arrays by a specified key. - **JinjaNode**: render Jinja2 templates against upstream data. ### Extraction Pipeline Nodes for structured data extraction from unstructured text: - **Extraction**: extract entities and relationships using AI. - **ConfidenceFilter**: filter extraction results by confidence threshold. - **EntityDeduplication**: merge duplicate entities. - **PromptingEntityExtraction** / **PromptingRelationExtraction** / **PromptingEntityDeduplication**: prompt-driven variants with full control over extraction behavior. - **Neo4jSink**: write extracted entities and relationships to a Neo4j graph database. ## Next Steps - Configure [triggers](/docs/workflow-triggers) to start your workflows - Learn about [execution and monitoring](/docs/workflow-execution) --- ## Workflow Triggers Source: https://aceteam.ai/docs/workflow-triggers Start workflows manually, on a schedule, via webhook, or from events. Every workflow begins with a trigger. The trigger determines when and how the workflow runs, and what initial data it receives. You must add exactly one trigger node to each workflow. There are four trigger types, each suited to different use cases. ## Manual Trigger A manual trigger lets you run the workflow on demand from the AceTeam interface. This is the simplest way to get started and is useful for workflows you want to run yourself, such as data processing tasks or one-time operations. To configure a manual trigger: 1. Add a **Manual Trigger** node to your canvas. 2. Optionally define input fields: these appear as a form when you click **Run**. For example, you can add a text field for a URL or a file upload field. 3. When you run the workflow, fill in the form and click **Execute**. Manual triggers are also helpful during development. You can test your workflow repeatedly with different inputs without setting up external infrastructure. ## Scheduled Trigger Scheduled triggers run your workflow automatically at regular intervals. They work like a cron job but are configured through a visual interface. To set up a schedule: 1. Add a **Scheduled Trigger** node. 2. Choose a frequency: every N minutes, hourly, daily, weekly, or monthly. 3. For daily and weekly schedules, pick the specific time and (for weekly) the day of the week. 4. Optionally set a timezone. All schedules default to UTC if no timezone is selected. Scheduled triggers do not pass any input data. If your workflow needs external data, combine the schedule with an HTTP node or database query as the first processing step. The platform shows the next scheduled run time on the workflow detail page. You can pause and resume schedules without deleting the trigger configuration. ## Webhook Trigger Webhook triggers start your workflow when an external system sends an HTTP POST request to a unique URL. This is how you connect AceTeam workflows to third-party services like Stripe, GitHub, Slack, or any tool that supports outgoing webhooks. To configure a webhook trigger: 1. Add a **Webhook Trigger** node. 2. The platform generates a unique URL automatically. Copy this URL and paste it into the external service's webhook settings. 3. Optionally define a **secret** for signature verification. When set, the platform validates the `X-Webhook-Signature` header on incoming requests and rejects unsigned or tampered payloads. 4. Choose whether to respond synchronously (the caller waits for the workflow to finish and receives the final output) or asynchronously (the caller receives a 202 Accepted response immediately). The full request body, headers, and query parameters are passed into the workflow as the trigger output. Downstream nodes can reference fields like `{{trigger.body.event_type}}` or `{{trigger.headers.authorization}}`. ## Event Trigger Event triggers react to things that happen inside the AceTeam platform itself. Use them to build automations that respond to platform activity. Supported events include: - **Agent message received**: Fires when a user sends a message to a specific agent. - **Workflow completed**: Fires when another workflow finishes execution. - **Team member invited**: Fires when a new member is added to your organization. - **File uploaded**: Fires when a file is uploaded to the knowledge base. To configure an event trigger: 1. Add an **Event Trigger** node. 2. Select the event type from the dropdown. 3. Optionally filter by specific resources (for example, only trigger on messages to a particular agent). The event payload is passed as the trigger output, containing details like the event type, timestamp, and associated resource data. ## Testing Triggers You can test any trigger from the editor without waiting for the real event: - **Manual**: Click **Run** and fill in sample data. - **Scheduled**: Click **Run Now** to simulate the next scheduled execution. - **Webhook**: Use the **Send Test Request** button, which sends a sample POST to your webhook URL. - **Event**: Click **Simulate Event** to fire the trigger with a sample event payload. Test runs appear in the execution history alongside real runs, marked with a "Test" label so you can tell them apart. ## Passing Data into Workflows Regardless of the trigger type, the trigger node's output is available to all downstream nodes through the `trigger` variable. Structure your workflow to reference trigger data explicitly, for example, `{{trigger.body.customer_email}}` for a webhook payload or `{{trigger.input.query}}` for a manual input field. This keeps your workflows readable and easy to debug. ## Next Steps - Learn about the [node types](/docs/workflow-nodes) you can connect after a trigger - Understand [workflow execution](/docs/workflow-execution) and monitoring --- ## Workflow Execution Source: https://aceteam.ai/docs/workflow-execution Run workflows, monitor progress, and handle errors. Once you have built a workflow in the editor, execution is where it comes to life. This page covers how to run workflows, monitor them in real time, inspect results, handle errors, and review past runs. ## Running a Workflow There are two ways a workflow executes: - **Manual run**: Open the workflow in the editor and click **Run** in the top bar. If the workflow has a manual trigger with input fields, a form appears for you to fill in before execution starts. - **Automatic run**: Workflows with scheduled, webhook, or event triggers execute automatically when their trigger condition is met. You do not need to have the editor open. When a workflow starts, the platform creates an **execution record** with a unique ID, a timestamp, and a status of "Running." ## Live Execution View While a workflow is running, the editor shows progress in real time. Nodes light up as they begin processing and change color to indicate their status: - **Gray**: Not yet reached. The node is waiting for upstream nodes to complete. - **Blue**: Currently executing. The node is processing its task. - **Green**: Completed successfully. The node finished and passed its output downstream. - **Red**: Failed. The node encountered an error and stopped. Edges animate to show data flowing between nodes. If your workflow has parallel branches, you will see multiple nodes executing at the same time. You can click on any node during execution to see its current state in the detail panel on the right side. ## Logs and Inspection After a node completes (or fails), click on it to inspect its execution details: - **Input**: The data the node received from upstream nodes. Displayed as formatted JSON. - **Output**: The data the node produced. For AI nodes, this includes the full model response. For HTTP nodes, this includes the status code, headers, and body. - **Duration**: How long the node took to execute, in milliseconds. - **Logs**: Any log messages the node generated during execution. Code nodes can write to logs using `console.log` (JavaScript) or `print` (Python). This level of visibility makes it straightforward to debug issues. If a downstream node produces unexpected results, trace back through the chain by clicking each node and comparing its input and output. ## Error Handling When a node fails, the workflow engine applies the following behavior: - **Node failure**: The node turns red and its error message is recorded. By default, a failed node stops execution of all downstream nodes in its branch. Other independent branches continue running. - **Automatic retries**: You can configure retry behavior on any node. Open the node settings and set the retry count (up to 3) and delay between retries (in seconds). The node will attempt execution again before marking itself as failed. - **Fallback outputs**: Some node types support a fallback output port. If the main execution fails and retries are exhausted, data routes through the fallback port instead, allowing you to build error-handling logic directly into your workflow (for example, sending a notification when an API call fails). - **Workflow-level failure**: If all active branches have failed or been blocked by upstream failures, the entire workflow execution is marked as "Failed." Error messages include details about what went wrong: HTTP status codes for failed API calls, model error responses for AI nodes, and stack traces for Code nodes. ## Execution History Every workflow maintains a history of past executions. Access it from the **History** tab on the workflow detail page. Each entry shows: - **Execution ID**: A unique identifier for the run. - **Status**: Completed, Failed, Canceled, or Running. - **Trigger type**: Whether the run was manual, scheduled, webhook, or event-driven. - **Start time and duration**: When the run began and how long it took. - **Node summary**: A count of successful, failed, and skipped nodes. Click on any historical execution to reload the editor with that run's data. You can inspect every node's input and output exactly as it was at the time of execution. This is useful for comparing results across runs, investigating intermittent failures, and auditing workflow behavior over time. You can filter the history by status, date range, or trigger type to find specific runs quickly. ## Canceling a Running Workflow If you need to stop a workflow mid-execution, click the **Cancel** button in the top bar while viewing a running execution. Cancellation is cooperative: nodes that are currently in progress will finish their current operation, but no new nodes will start. The execution status changes to "Canceled." Canceled runs appear in the execution history with their partial results intact. You can inspect which nodes completed before the cancellation took effect. ## Next Steps - Review [workflow nodes](/docs/workflow-nodes) to expand what your automations can do - Set up [triggers](/docs/workflow-triggers) for automatic execution - Return to the [workflow editor](/docs/workflow-editor) to refine your workflows --- ## Human Approval Steps in Workflows Source: https://aceteam.ai/docs/human-approvals Add a human-in-the-loop approval gate to a workflow, then review, approve, or deny pending items from the review queue or over MCP. Some workflows should not run to completion on their own. Before a step that sends money, emails a customer, or writes to a system of record, you may want a person to look at what the workflow produced and decide whether it continues. The Human Approval node is a human-in-the-loop checkpoint you can drop anywhere in a flow: the run pauses, a pending item appears in your review queue, and the run resumes only after a human approves or denies it. ## How the Approval Gate Works The Human Approval node takes a single input, a `payload` (any JSON value, such as a draft, a computed result, or upstream context), and passes that same `payload` through unchanged on its output once the run is approved. Downstream nodes receive exactly what the reviewer saw. When the run first reaches the node, it does two things: 1. It writes a pending review item into your organization's review queue, carrying the node's `payload` so a reviewer can see what they are approving. 2. It yields, pausing the run. The workflow status becomes yielded, and no downstream node executes yet. The run stays paused until a human acts: - **Approve** resumes the run and passes the `payload` through to the next node. - **Deny** resumes the run, which then stops. In the current version the node has no separate deny branch, so a denial ends the run in `ERROR` status with the denial reason recorded in the run's error, not `CANCELLED`. The reviewer is anyone in your organization who can reach the Review and Approval queue. There is no per-node approver assignment. ## Node Configuration The Human Approval node exposes these parameters in the workflow editor: - **Prompt** (`label`): The text shown to the reviewer to explain what they are checking. Defaults to "Approval required". - **Timeout (seconds)** (`timeout_seconds`): How long to wait for a human before auto-resolving. `0` (the default) waits indefinitely. - **Timeout action** (`timeout_action`): What happens once the timeout elapses with no human action, either `deny` (the default) or `approve`. - **Risk tag** (`risk_tag`): An optional reversibility signal shown to the reviewer, using the same vocabulary as tool-call approvals: `read_only`, `reversible_write`, or `one_way_door`. Leave it blank for no signal. A note on the timeout: the deadline is only evaluated the next time the run is naturally resumed. Nothing wakes a paused run on its own when the deadline passes, so a timeout resolves a stale wait rather than acting as a precise scheduled deadline. ## Adding an Approval Step 1. Open the workflow you want to gate in the editor. 2. Add a **Human Approval** node on the path where the checkpoint belongs, between the step that produces the value to review and the step that acts on it. 3. Connect the upstream node's output into the Human Approval node's `payload` input. 4. Connect the Human Approval node's `payload` output into the downstream node that should run only after approval. 5. Set the **Prompt** so the reviewer understands what they are approving. 6. Optionally set **Timeout (seconds)** and **Timeout action**, and a **Risk tag** to signal how reversible the gated action is. 7. Save the workflow. The node type is `HumanApproval` and it belongs to the `control_flow` category, so you can also find it with `flow_search_node_types` and inspect its ports with `flow_node_type_details` when building a graph over MCP. ## Reviewing Pending Items Pending approval items appear in your organization's Review and Approval queue in the web console (the `/reviews` surface). Each item shows the workflow, the run, the node, and the `payload` awaiting a decision. You can approve or deny any node type from the web console. You can also review and act on Human Approval items over MCP with these tools: - `workflow_review_list` lists review items. By default it returns only pending items; pass `all=true` to also see reviewed, skipped, and aborted rows. Each row carries the `workflow_run_id`, `node_type`, `node_id`, `round`, `status`, `created_at`, and, when the author set one, the `risk_tag`. - `workflow_review_get` returns the full detail of one review item by `review_id`, including the payload under review. Call it before deciding, to see what you would be approving. - `workflow_review_approve` approves a pending item by `review_id` and resumes the yielded run. An optional `reply` is recorded alongside the approval. - `workflow_review_deny` denies a pending item by `review_id`. An optional `reply` is recorded as the denial reason and surfaced in the run's error message. To act on an item over MCP: 1. Call `workflow_review_list` to find the pending item and its `review_id`. 2. Call `workflow_review_get` with that `review_id` to read the payload under review. 3. Call `workflow_review_approve` or `workflow_review_deny` with the `review_id`, and an optional `reply`. Approve and deny use the same resume path as the web console's buttons, so a run resumes exactly as it would from a human click. ### What the MCP Review Tools Can and Cannot Act On The review queue also holds review rounds for other node types, such as `AgenticDrafting` and `AgenticClassification`, whose review shape is specific to those nodes rather than a plain approve or deny decision. `workflow_review_list` and `workflow_review_get` show those items for visibility, but only `HumanApproval` items can be approved or denied through `workflow_review_approve` and `workflow_review_deny`. For any other node type, act on the item from the web console's Review and Approval queue instead. ## Not the Same as the Agent Approvals Queue The `approvals_list`, `approvals_get`, and `approvals_cancel` tools operate on a different queue: the agent tool-call approval queue, which holds gated agent actions (for example a queued `email_send`), not workflow Human Approval reviews. That surface is read-only plus self-cancel: `approvals_cancel` can only withdraw one of the caller's own still-pending actions, and there is no approve or deny on that surface by design. Use the workflow review tools above, or the web console, for Human Approval steps in a flow. A tool classified `one_way_door` on the agent tool-call approval queue is meant to always route through this gate before it runs. If a tool carries that classification but has no gate actually wired to it, the platform refuses the call outright for an agent rather than letting it run unheld: a human working in the web console keeps normal access. This is a fail-closed backstop, not a policy setting, so it cannot be waived or auto-approved around. ## Next Steps - Learn about the [node types](/docs/workflow-nodes) you can gate with an approval step - Configure [triggers](/docs/workflow-triggers) that start your workflows --- ## Schedule Shell Commands, Workflows, Agent Prompts, and Reminders Source: https://aceteam.ai/docs/scheduled-jobs-reminders Create recurring or one-shot scheduled jobs that run a shell command on your own node, trigger a workflow, run an inline agent prompt, or deliver a static reminder, and inspect each run's history. This page covers the standalone scheduled-jobs surface: `cron_create`, `cron_list`, `cron_update`, `cron_delete`, `cron_history`, and `reminder_create`. It is a separate mechanism from the workflow-level scheduled trigger described in [Workflow Triggers](/docs/workflow-triggers), which starts a specific workflow on its own schedule from inside the workflow itself. This surface instead manages a standalone job row that can point at a shell command, a workflow, an inline agent prompt, or a static reminder message, and fires it on a cron schedule or once at a chosen time. ## The Four Job Targets A job row can carry any of `command`, `workflow_id`, `agent_prompt`, or a reminder marker (set via `reminder_create`, described below). When more than one is set on the same row, the platform's executor picks exactly one target per fire, in this order of precedence: reminder, then workflow, then command, then agent prompt. - **`workflow_id`**: triggers a run of that workflow on each fire, the same way any other trigger would. - **`command`**: runs on your organization's own Citadel node, never on the shared coordinator host. This is a tenant-safety decision: the command is organization-supplied text, so it cannot be allowed to run as arbitrary shell on infrastructure other tenants also run on. If your organization has no Citadel node connected, the run is recorded as failed with a clear message. - **`agent_prompt`**: runs your organization's Ace agent inline as one user turn with its full toolset, and pushes the final answer to your own registered devices through the same delivery path `notify`'s push channel uses. This is the lightweight "check something, decide, page me" path with no pre-authored workflow needed. The turn is metered against your organization's ACET balance like any other hosted agent invocation, and an insufficient-balance denial is recorded as a clean failed run rather than a crash. - **Reminder** (via `reminder_create`): delivers a stored, static message with no agent turn and no model call at fire time. See below. ## Scheduling: Recurring or One-shot `cron_create` and `reminder_create` both take exactly one of a recurring schedule or a one-shot fire time: - `schedule`: a five-field cron expression, `minute hour day_of_month month day_of_week` (for example `30 9 * * 1-5` for weekdays at 9:30). - `run_once_at` (on `cron_create`) or `run_at` (on `reminder_create`): an ISO 8601 timestamp in the future. A one-shot job disables itself automatically after it fires once, so it never fires a second time on its own; re-enabling it with `cron_update(enabled=true)` re-arms and re-fires it. ## Creating a Job `cron_create` takes a `name`, exactly one of `schedule` or `run_once_at`, and at least one of `command`, `workflow_id`, or `agent_prompt`. `description` and `enabled` (default `true`) are optional. `cron_update` changes an existing job by ID; only the fields you pass are changed, and `schedule`/`run_once_at` remain mutually exclusive on the update as well. `cron_delete` removes a job outright. ## Static Reminders `reminder_create` is a distinct, lightweight tool for scheduling a message to yourself with no agent involved at all: "text me at 9am to submit the report," or a recurring nudge to check a queue. It is backed by the same underlying job table as `cron_create`, storing its `message`, `channel`, and `severity` behind a marker in the row rather than needing its own schema, but it is not the same as an `agent_prompt` job: nothing is generated at fire time, only the exact stored `message` is delivered, through the same delivery path `notify` uses, defaulting to the `push` channel (also supports `sms`, `whatsapp`, `wechat`, and `telegram`). Recurring reminders are rate-limited per severity, the same protection `notify` applies, so a mistakenly tight schedule cannot flood your phone; `critical` severity is exempt from the limit. A reminder shows up in `cron_list` alongside ordinary jobs and is canceled with the same `cron_delete`. `cron_update` can retime, rename, or enable and disable a reminder, but rejects an attempt to change its `command`, `workflow_id`, or `agent_prompt` fields, since converting a reminder's target column into something else is exactly the kind of mis-route the underlying storage design guards against. ## Listing and Inspecting Runs `cron_list` returns every scheduled job for your organization as a table: ID, name, schedule, target, and whether it is enabled. `cron_history` returns a job's execution history (started time, status, duration, and a truncated output snippet) for a given `job_id`, capped at 50 rows per call. ## What Happens When a Run Fails A failed `command` or `workflow` run is logged and its status is recorded on the run row, but nobody is paged about it; there is currently no general per-job failure notification for those two targets. The `agent_prompt` path is the one delivery-bearing target: its whole point is to push a result (or a clean failure note, if the agent turn could not be billed) to your devices. Reminders are delivery by design, so a reminder that is rate-limited still records a clean successful run whose output notes that nothing was actually delivered that time. ## Typical Flow 1. Decide which target fits: a workflow you already built, a shell command that only needs to run on your own hardware, an inline agent decision with no pre-built workflow, or a plain reminder with no agent at all. 2. `cron_create` (or `reminder_create` for the plain-message case) with either a recurring `schedule` or a one-shot fire time. 3. `cron_list` to confirm the job is registered and enabled. 4. `cron_history` after it has fired at least once, to confirm it ran and to read its recorded output or failure. --- ## Connect a Database, Publish a Flow as an Endpoint, and Receive Webhooks Source: https://aceteam.ai/docs/database-connections-endpoints Give a Flow a database credential to dial directly, publish it as a callable HTTP endpoint or a public run receipt, and register a webhook to receive platform events. Three related tools let a Flow reach past the AceTeam platform: a database connection a Flow's SQL nodes can dial directly, a publish step that turns a finished Flow into something a stranger can call or read, and a webhook subscription that lets an outside system hear about platform events as they happen. ## Database Connections A Flow's `SQLSelect` and `SQLInsert` nodes take a `connection_id` naming a stored database credential. Before that credential can exist, someone has to create it. `connection_create` stores a PostgreSQL connection (host, port, database, user, password) for your organization. It tests reachability first, over the exact same SQLAlchemy code path `SQLSelect` and `SQLInsert` dial at run time, not a separate check, and refuses to save a connection it cannot reach. The password is encrypted at rest with a key derived per organization, and is never returned, logged, or echoed back by any tool, including a later read of the same connection. `connection_list` returns the connection ids your Flows can reference, without the password. `connection_test` re-runs the same reachability check against an already stored connection, so you can confirm a Flow that references it will actually be able to connect before running it. ## Publishing a Flow as a Callable Endpoint `flow_publish_endpoint` turns a Flow into something callable from outside an authenticated AceTeam session, at one of three visibility levels: - `private` and `org` do not mint anything new. They point back at the existing authenticated run route, scoped to your organization by its API key, for a caller who already holds one. - `link` mints a signed, unguessable URL with no authentication header at all: post the Flow's input fields as JSON, get its output fields back as JSON. Anyone holding the link can trigger the Flow, running as your organization, so publishing a Flow this way runs the same node type safety gate used elsewhere on the platform before the link is minted. Calling `flow_publish_endpoint` again on an already published Flow returns the existing link unchanged; pass `rotate` to mint a fresh token and invalidate the old one. `flow_unpublish_endpoint` revokes the endpoint; unpublishing an already unpublished Flow is a clean no-op. ## Publishing a Run as a Public Receipt `flow_run_publish_share` is a different, narrower kind of publish: instead of letting a stranger trigger a Flow, it lets a stranger read the record of one run that already happened, with no access to the Flow itself or your organization. It refuses to publish a run whose Flow contains a side effecting node type, such as Shell, Email, or an API call, anywhere in its graph including inside a nested sub flow, and it refuses a run that has not finished yet. What the public read exposes is a fixed set of fields only: the run's status and timing, each node's status and timing plus whether it errored (never the error text itself), aggregate cost and token usage, and the full provenance receipt chain. Run and node input, output, and error text are never exposed. Calling `flow_run_publish_share` again on an already published run returns the existing link; pass `rotate` to mint a fresh one, or call `flow_run_unpublish_share` to revoke it outright. ## Webhooks Where a published Flow endpoint lets an outside system call into AceTeam, a webhook lets AceTeam call out. `create_webhook` registers a URL to receive an HTTP POST for a chosen set of platform event types, such as a completed or failed job, a node coming online or offline, a service starting or stopping, a sandbox suspending or resuming, a low ACET balance, a hosted surface submission, or review activity. Registering a webhook returns an HMAC secret shown exactly once. Use it to verify the `X-AceTeam-Signature` header on every incoming request, since the secret itself is never shown or returned again. `list_webhooks` lists your organization's registered webhooks with their URLs and subscribed event types, without exposing any secret.