34 open source tools compared. Sorted by stars. Scroll down for our analysis.
See our ranked picks: Best Open Source Agent Runtimes & Sandboxes
By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.
| Tool | Stars | Velocity | Score |
|---|---|---|---|
cc-switch A cross-platform desktop All-in-One assistant tool for Claude Code, Codex, OpenCode, openclaw & Gemini CLI. | 136.1k | +2791/wk | 91 |
deer-flow An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours. | 82.9k | +308/wk | 95 |
lobehub 🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team. | 82.8k | +279/wk | 85 |
mempalace The best-benchmarked open-source AI memory system. And it's free. | 59.3k | +114/wk | 93 |
mempalace The highest-scoring AI memory system ever benchmarked. And it's free. | 59.1k | - | 87 |
cherry-studio AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs | 52.1k | +184/wk | 87 |
QwenPaw Your Personal AI Assistant; easy to install, deploy on your own machine or on the cloud; supports multiple chat apps with easily extensible capabilities. | 35.3k | +179/wk | 92 |
supermemory Memory engine and app that is extremely fast, scalable. The Memory API for the AI era. | 30.0k | - | 92 |
NemoClaw Run OpenClaw more securely inside NVIDIA OpenShell with managed inference | 22.5k | +45/wk | 97 |
hermes-webui Hermes WebUI: The best way to use Hermes Agent from the web or from your phone! | 18.4k | - | 80 |
BrowserOS 🌐 The open-source Agentic browser; alternative to ChatGPT Atlas, Perplexity Comet, Dia. | 13.7k | +35/wk | 82 |
ironclaw IronClaw is an Agent OS focused on privacy, security and extensibility | 12.6k | - | 84 |
bifrost Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS. | 8.1k | - | 78 |
Steel Browser 🔥 Open Source Browser API for AI Agents & Apps. Steel Browser is a batteries-included browser sandbox that lets you automate the web without worrying about infrastructure. | 7.7k | - | 62 |
engram Persistent memory system for AI coding agents: agent-agnostic Go binary with SQLite + FTS5, MCP server, HTTP API, and CLI. | 6.8k | +97/wk | 80 |
agent-governance-toolkit AI Agent Governance Toolkit: policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10. | 6.3k | - | 78 |
byterover-cli ByteRover CLI (brv) - The portable memory layer for autonomous coding agents (formerly Cipher) | 5.0k | +1/wk | 49 |
kungfu Continuity for Agent Work | 4.5k | +8/wk | 80 |
archestra Enterprise AI Platform with guardrails, MCP registry, gateway & orchestrator | 4.3k | +20/wk | 68 |
KiroCrew A persistent workspace for development work that self-improves and continues beyond one session. | 4.1k | +159/wk | 78 |
PilotDeck Task-oriented AI Agent productivity platform | 4.0k | +7/wk | 70 |
openclaw-control-center Turn OpenClaw from a black box into a local control center you can see, trust, and control. | 4.0k | +1/wk | 66 |
mirage A Unified Virtual Filesystem For AI Agents | 3.6k | - | 74 |
AgentENV AgentENV (AENV) is a distributed platform for running agent environments at scale. | 3.5k | +48/wk | 76 |
Qclaw 不用命令行,小白也能轻松玩转 OpenClaw | 2.8k | - | 61 |
cindy Consider it done. The open-source AI agent that works out of the box · 想到,就能做到。开源、开箱即用的 AI Agent。 | 2.7k | - | 70 |
boxlite Sandboxes for every agent: embeddable, stateful, with snapshots and hardware isolation. | 2.3k | +10/wk | 76 |
codex-console codex-console 是一个集成化控制台项目,支持任务管理、批量处理、数据导出、自动上传、日志查看与打包支持。 | 2.2k | +5/wk | 68 |
gitlab-mcp First gitlab mcp for you | 2.0k | - | 68 |
llm-space A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first, cloud-ready for managed agents. | 1.9k | +23/wk | 68 |
mcp-brasil MCP Server para 41 APIs públicas brasileiras | 1.8k | +11/wk | 72 |
OpenSquirrel For people who get distracted by agents. A native Rust/GPUI control plane for running Claude Code, Codex, Cursor, and OpenCode side by side, because if you're going to be squirrely, you might as well optimize for it. | 1.4k | - | 59 |
commonly Open-source room for humans and cross-vendor AI agents. Every agent gets its own name, memory, skills, and workstation. Any runtime, your infra, no per-agent fees. | 1.4k | +8/wk | 72 |
vm0 the easiest way to run natural language-described workflows automatically | 1.2k | - | 59 |
Stay ahead of the category
New tools and momentum shifts, every Wednesday.
Cc-switch wraps them into a single Tauri-based GUI. Cross-platform, open source, and free. Consider it a launcher that lets you switch between agents without context-switching between terminals. Setup is straightforward: download the app, configure your API keys, and pick which agents you want active. It doesn't add intelligence on top of the agents themselves. It's a convenience layer. The value is entirely in the unified interface and the ability to compare agent outputs side-by-side. Solo developers who already use multiple coding agents will get the most out of this. Teams probably don't need it since most teams standardize on one agent. If you only use one coding agent, there's nothing here for you. The catch: it's a wrapper, not a product. If the underlying agents change their CLI interfaces (which they do, frequently), cc-switch breaks until someone updates the integration. You're adding a dependency on a third-party GUI for tools that already work fine in a terminal.
DeerFlow is ByteDance's open source answer to Manus and ChatGPT's agent mode: an agent harness built to work a task for minutes or hours, not just answer a prompt. It ships the whole stack (backend, web UI, sandboxed execution, chat gateways for Slack, Telegram, Discord, and more) and gives the agent a real filesystem, persistent memory, loadable skills, and sub-agents to delegate to. MIT licensed, completely free. Self-hosting is real work. Python, Node, and nginx, an interactive setup wizard, Docker as the recommended path, and a hardware floor around 4 vCPU and 8 GB of RAM just to evaluate it. It's model-agnostic through any OpenAI-compatible API. The README is refreshingly blunt that improper deployment introduces security risks: this thing runs code, so it's designed to stay bound to localhost unless you put auth and network isolation in front of it. Use it if you want long-horizon agents on your own infrastructure with your choice of models. Solo and small teams: free, plus tokens. If you just want agent orchestration as a library, LangGraph (which DeerFlow builds on) or CrewAI is the lighter path, and OpenHands is the closest open comparison. The catch: 2.0 is a ground-up rewrite that shares no code with v1, so the project is popular but the current codebase is young, with a long open-issue tail. And an agent with shell access is a loaded tool. Mount credentials into its sandboxes carelessly and you'll regret it.
LobeHub is the rebrand of LobeChat, and it grew up. What started as a slick self-hosted ChatGPT interface is now an agent operations platform: you build a team of AI agents, give them schedules and skills, and let them run tasks around the clock. It plugs into OpenAI, Claude, Gemini, and local models, with thousands of MCP plugins for tools and data. Open source under LobeHub's own community license, free to self-host, with a hosted cloud tier if you would rather not run it. Self-hosting is a Docker job, and it is genuinely one-click on Vercel, Zeabur, or Sealos when all you want is the chat experience. The interface is the easy part. You bring your own model API keys, so the real cost is whatever OpenAI or Anthropic charges per token, not the software. Running the full agent-operations layer with scheduling, agent groups, and persistent memory is more involved than the basic chat deployment. Solo builders should self-host: free, trivial setup, and a private front end for every model you use. The cloud tier meters credits rather than seats, which changes the math. Free gives you 500,000 credits a month, 10 MB of storage, and a short model list. Starter is $12.90 a month, or $9.90 if you pay yearly, for 5 million credits and access to Claude Sonnet 5. Premium is $24.90 monthly or $19.90 yearly for 15 million, and Ultimate is $49.90 or $39.90 for 35 million. Small teams that would rather buy credits than run Docker land on Starter or Premium. Enterprise is quoted for private deployment. The catch is the license. It is not plain MIT, and any organization standardizing on this internally needs to read the community license terms before it becomes load-bearing. The 24/7 agent-team story is also new, and how it holds up under real production load is still unproven.
MemPalace stores your AI conversation history verbatim and searches it semantically. Every Claude session, every project file, indexed locally. The structure is a metaphor: projects become wings, topics become rooms, so you can scope searches instead of querying a flat blob. It publishes 96.6% recall on LongMemEval with no LLM in the loop, and the benchmarks are reproducible from the repo. Install is pip plus pointing it at a directory. ChromaDB is the default backend, embeddings run on CPU with a 300MB model, no API key required. The MCP server exposes 29 tools so Claude Code can read and write the palace directly during a session. Solo developers using Claude Code heavily: install it. The 'wake-up' command that loads relevant context for a new session is the pitch and it works. Small teams: each engineer runs their own palace, there is no shared knowledge layer yet. The catch: it's about two weeks old. The benchmarks are real but the operational track record is not. Breaking changes will happen, and fast-growing projects attract impostor domains. The README has a scam alert for a reason.
MemPalace gives your AI assistant a long-term memory you actually own. It stores conversations and documents verbatim, nothing summarized away, then lets the assistant pull back the exact relevant pieces with semantic search. MIT licensed, free, no API key required, organized around a memory-palace metaphor: wings for people and projects, rooms for topics, drawers holding the original content. It started life strictly local, and that's still the default: a Python install, a pluggable vector backend (ChromaDB by default; Qdrant, pgvector, Milvus, and SQLite all work), and an MCP server with auto-save hooks for Claude Code. It has since grown past one machine. Recent releases added a team server with TLS and authentication, an HTTP transport for the MCP server, and a coordination layer that lets agents on different machines hand work to each other. Solo developers who want their assistant to remember context across sessions without shipping every conversation to a vendor: this is one of the strongest local options going. Small teams can now share one memory server instead of syncing folders. The closest comparison, supermemory, leans on a hosted API; MemPalace's pitch is that everything stays yours. The catch is pace. It ships releases constantly, and the feature set moves month to month. Pin your version if you depend on it.
Cherry Studio puts every LLM behind one desktop window: chat with OpenAI, Claude, Gemini, or a local Ollama model from the same app, with hundreds of preconfigured assistants, knowledge bases, document processing, and MCP support for calling external tools. Download it for Windows, Mac, or Linux and it just runs. No server, no Docker. The Community Edition is AGPL-3.0, and unusually for dual-licensed projects, there's no seat cap or revenue threshold hiding in the license file: commercial use is allowed outright as long as you comply with AGPL. The paid commercial license exists purely to buy an exemption from AGPL obligations, and a separate Enterprise Edition adds private deployment, an admin console, RBAC, and quota management at quote-only pricing. Solo and small teams: free, plus whatever API keys you plug in. Open WebUI and LibreChat are the browser-based comparisons; Jan and Chatbox are the desktop ones. Cherry Studio's edge is breadth: assistants, agents, knowledge bases, and MCP in one polished client. The catch: the enterprise side is oriented to the Chinese market (Alipay and bank transfer, quote-based sales), the license terms have been revised repeatedly over the project's life, and the issue tracker runs long. Read the AGPL before you fork it into a product.
QwenPaw is a personal AI assistant you run yourself instead of renting from OpenAI. It ships with a local runtime so it works with no API key out of the box, and it also plugs into Ollama, LM Studio, and a dozen-plus cloud providers if you want bigger models. The hook is reach: it talks to you through Discord, Telegram, Lark, WeChat, DingTalk, even iMessage, and you extend what it can do with skills. It is open source under Apache-2.0, built by the team behind AgentScope, Alibaba's multi-agent framework. Self-hosting is the default here, not an afterthought. There is Docker support and a one-click path to deploy on Alibaba Cloud if you would rather not run it at home, in which case you pay for the cloud, not the software. It takes its own security seriously for a personal tool: a kernel-level sandbox, a Tool Guard, and a File Guard sit between the model and your machine, which matters once an assistant can run code and touch your files. For a solo developer or a tinkerer who wants an assistant that lives in their own chat apps and on their own hardware, this is one of the more complete self-hosted options going, and it costs nothing. Small teams can share an instance. There is no real large-team story here; it is a personal workstation, not a company-wide deployment, and that is fine. The catch is gravity. It is deep in the Alibaba and Qwen ecosystem, the docs are heavily multi-language, and a lot of the built-in channels (WeChat, DingTalk, Lark) point at a Chinese user base. None of that is a flaw, but if you expected a Western-defaults, English-first assistant, calibrate before you install.
Supermemory is a memory and context engine for AI apps. It extracts facts from conversations and documents, keeps a profile of each user, and runs hybrid search that blends RAG with personalized recall, all behind one API. The company reports top scores on the major memory benchmarks. The main repository is MIT, and plugins exist for Claude Code, Cursor and Codex.
Running it locally is one command, npx supermemory local, which starts a single binary with the same API as the hosted platform. It works fully offline with a local model through Ollama, and your data stays on your machine. The local server is single-user and licensed for up to 10,000 documents, though, and the connectors for Google Drive, Gmail, Notion and OneDrive, plus the Supermemory MCP, only exist on the hosted platform.
Hosted pricing is published and credit-based. Free includes $5 of credits a month and pauses when they run out. Pro is $19/mo with pay-as-you-go top-ups, Max is $100/mo, and Scale at $399/mo adds S3 and web crawler connectors, unlimited seats and a self-hosted option. Teams building production memory will land on Pro or Scale.
The catch is maturity. The local server is still version 0.0.x, and one August release was an emergency patch after an upgrade silently wiped search vectors. The consumer app and browser extension were retired in September too. Build on the API, pin versions, and back up your data.
NemoClaw runs OpenClaw (the open source coding agent) inside NVIDIA's OpenShell sandbox with managed inference, solving the real security risk of agents executing arbitrary code on your machine. Your agent gets GPU-accelerated model inference through NVIDIA's infrastructure while staying sandboxed. This is NVIDIA saying 'run your coding agents on our hardware, securely.' You get the performance of NVIDIA GPUs for inference without managing the infrastructure yourself. The sandbox prevents the agent from doing anything destructive to your system. Apache 2.0 licensed. The catch: this ties you to NVIDIA's ecosystem. You need NVIDIA hardware or their cloud infrastructure, no running this on Apple Silicon or AMD GPUs. It's OpenClaw-specific, so Claude Code and Cursor users are out. And 'managed inference' is a gateway to NVIDIA's paid compute. The tool is free but the GPU time may not be.
Hermes WebUI is a browser frontend for Hermes Agent, a self-hosted autonomous AI agent that holds memory across sessions, runs scheduled jobs, and integrates with messaging platforms. Free and MIT-licensed. Setup is moderate. You bring your own LLM API key (OpenAI, Anthropic, Google, DeepSeek, OpenRouter, others) and run the agent plus WebUI on your own hardware or VPS. Once running, the agent persists conversation context, learns from interactions, and can be triggered on a schedule. The web UI mirrors the CLI experience without locking you out when you close the terminal. For solo developers and small teams who want an AI agent that isn't tied to ChatGPT or Claude.ai, this is a real option. Your conversations, your memory, your hardware. The cost is your LLM API bill, which can climb fast if the agent is making frequent calls. Solo: probably $10 to $50 per month in API spend depending on usage. The catch is that "autonomous AI agent" is doing a lot of work in the description. These systems still hallucinate, still drift, still need supervision. Don't wire it into anything destructive without guardrails.
BrowserOS is a Chromium fork built for AI agents to drive. The neo variant runs as a second browser where Claude Code, Cursor, or any MCP-compatible agent does web tasks in its own tabs using your logged-in accounts, with every session recorded and replayable as scrubbable video. The full browser variant is a daily driver with an agent built in. AGPL, free, bring your own model keys. Setup is signed installers for macOS and Windows, with Linux on the full browser. There's no BrowserOS subscription; cost is whatever your model provider charges, and local models via Ollama bring that to zero. Building from source is a Chromium build, so plan on hours, not minutes. This is for developers already using coding agents who want them doing authenticated web work, pulling data from dashboards, filling forms, without handing credentials to a cloud service. The proprietary comparisons are ChatGPT Atlas, Perplexity Comet, and Dia. The catch is trust and blast radius. You're giving a young third-party Chromium fork your live sessions, and an agent holding your cookies can act as you. Chromium forks also historically lag upstream on security patches. Keep it away from anything you can't afford to have clicked.
IronClaw is a personal AI assistant that runs entirely on your own machine. Built by NEAR AI, it's what you reach for when you want an always-on agent that reads your email, runs scheduled jobs, and answers from Telegram or Slack, but you don't want any of that data leaving your control. Everything is stored locally and encrypted. It's dual-licensed Apache 2.0 and MIT, completely free and open source.
It ships as a single Rust binary, which is the whole pitch: native speed, memory safety, nothing extra to babysit. Install is a shell script or Homebrew, then ironclaw onboard wires up your LLM provider. It leans on Postgres for persistence rather than SQLite, and runs untrusted tools inside a WASM sandbox or Docker, so the security story is built in, not bolted on. You bring your own model API keys, and you'll want Postgres running somewhere.
This is a young project, a Rust reimplementation inspired by OpenClaw, so treat it as early but serious. Solo devs and privacy-minded tinkerers: this is the fun one, a local agent you actually control. Small teams: usable for internal automation if someone's comfortable with Rust and Postgres. Large teams: watch it, don't bet a production workflow on it yet.
The catch is that "Agent OS" is carrying a lot of weight in the description. You're running an early-stage framework, not a finished product, and the work of wiring up providers, keys, and a database is on you. The privacy guarantee is only as good as the setup you build around it.
Bifrost is an open source gateway that puts one OpenAI-compatible API in front of 20+ LLM providers (OpenAI, Anthropic, Bedrock, Vertex, and others). Same idea as LiteLLM, written in Go, with the team claiming significantly higher throughput at concurrency. Apache 2.0 and free to run yourself. Setup is as easy as it gets in this category: 'npx -y @maximhq/bifrost' to try, or Docker and a config file for real use. It has automatic failover, load balancing, semantic caching, MCP integration, plus governance pieces like per-team budgets and rate limits. A web UI handles config, so you don't have to live in YAML. For solo developers and small teams routing AI calls across providers, this is a direct LiteLLM alternative with less Python overhead. Larger teams comparing both will care about throughput claims under real load, verify on your traffic not the benchmark. Open core, with an enterprise tier for clustering, adaptive load balancing, and custom plugins. The catch: it's newer and less battle-tested than LiteLLM, and headline benchmarks rarely match production. If LiteLLM is already running cleanly, the migration story has to clear a real bar. If you're picking now, Bifrost's setup speed is a real win.
Steel Browser is browser infrastructure for AI agents: a Chrome instance behind an API that an agent drives, with sessions that keep cookies and local storage alive between requests. It is the self-hosted answer to paying by the browser hour, and it is Apache-2.0. Docker is the fast path, and it also deploys to Railway or Render or runs on bare Node with Chrome. Control is Puppeteer and CDP, so everything you already know about driving Chrome still applies. What it saves you is assembly: stealth plugins and fingerprint management, proxy chain support for IP rotation, Chrome extension loading, and endpoints that turn a page straight into markdown, a screenshot, or a PDF. Solo developers and small teams building agents that browse: self-host it and the marginal cost of a session becomes your own compute. Medium teams: still self-host, but budget real operational time, because headless Chrome at concurrency is memory hungry and crashed sessions are normal rather than exceptional. Large teams running a fleet: compare honestly against Steel Cloud or Browserbase, because managed browser services exist precisely because scaling this is miserable. The catch: releases are still tagged beta and arrive months apart even though commits keep landing. The bigger one is that self-hosting means you personally own the anti-bot arms race. The stealth plugins are in the box, but keeping them effective against sites that actively fight automation is continuous work, and absorbing that work is most of what the hosted providers are actually charging for.
Engram gives it persistent memory. It's a Go binary with SQLite and full-text search that any AI agent can read and write to, so context survives across sessions. It works via MCP server, HTTP API, or CLI, meaning it's agent-agnostic. Claude Code, Codex, OpenClaw, or anything else that speaks HTTP can use it. Your agent writes memories during a session and reads them back next time. Full-text search (FTS5) means it retrieves relevant context, not just raw dumps. MIT licensed, Go. The catch: persistent memory is only useful if the agent writes good memories. Garbage in, garbage out. If the agent stores irrelevant context, it pollutes future sessions. SQLite is great for single-user but won't scale to a team sharing one memory store. And the MCP protocol is still young; not every agent supports it natively.
Agent Governance Toolkit puts a policy check between your AI agent and every action it takes. Tool calls, resource access, and agent-to-agent messages all get evaluated against rules you write before anything runs. The distinction that matters: this is application-layer enforcement, not prompt instructions asking a model to behave, so a jailbreak in the conversation does not talk its way past it. MIT, maintained by Microsoft, free. It is framework-agnostic on purpose, with adapters for LangChain, CrewAI, AutoGen, AWS Bedrock, Google ADK, and Azure AI among others, plus SDKs for Python, TypeScript, Rust, Go, and .NET. You get a CLI, a governance dashboard, and published specs mapped against the OWASP Agentic Top 10. Installation is a package manager command; the actual work is writing policy that reflects what your agents should be allowed to do, which nobody can do for you. Someone shipping a weekend agent can skip this. Anyone putting an autonomous agent in front of production systems or customer data needs this or an equivalent, and there is no free-versus-paid decision to make because Microsoft ships the whole thing MIT with no commercial gate. The catch is scope, and it is narrower than the marketing suggests. This governs what an agent does, not what a model says, so prompt injection that produces a bad answer rather than a bad action passes straight through and you still need a content safety layer beside it. It is also public preview, with breaking changes expected before GA.
ByteRover adds a persistent memory layer that travels with you. It works as a CLI tool that sits alongside Claude Code, Codex, or any agent that reads context files. It's a portable brain for your coding assistant.
Install it globally, run brv init in your project, and it creates a structured memory store. The agent can read and write to it during sessions, building up project knowledge over time. It stores things like architecture decisions, coding conventions, and task history. The data lives on your machine in JSON files.
This solves a real problem for developers who spend the first 5 minutes of every AI session re-explaining their project. Solo developers and small teams get the most value. The memory is project-scoped, so each repo gets its own context.
The catch: you're trusting a third-party tool to manage context that feeds directly into your AI agent. If the memory format drifts from what agents expect, or if the project goes unmaintained, you've got stale context files that might do more harm than good. And Claude Code already has its own CLAUDE.md convention for project context, so the overlap is real.
Kungfu tackles the thing that quietly breaks long agent sessions: context loss on a handoff. It's a continuity layer that preserves task state and context so agents like Codex, Claude, OpenCode, and Amp can pause, hand off, and resume without forgetting what they were doing. The project wraps this in unusually formal governance, cryptographic "Release Passports" and a written qualification spec. It's a real, deep monorepo spanning TypeScript, Rust, and C++, with a serious commit history behind it. This is not a thin repo coasting on a good README, the code is there and the engineering is deliberate. For teams running agents on multi-step work, the promise is fewer dropped threads and cleaner resumption. Solo or team, the license cost is zero. The catch is timing. Kungfu is alpha. The public packaging hasn't shipped ("v4 coming soon"), so today you're building from source, not running an install command. The ideas are strong and the foundation looks solid, but this is one to watch and test, not to put in front of production work yet.
Archestra sits between your AI agents and your company's data, and tries to make that connection safe enough for a real enterprise. It is an open-source control plane: an LLM gateway that fronts any model provider, a registry and gateway for MCP servers, an agent orchestrator, and a layer of guardrails (SSO, RBAC, sandboxed code execution, prompt-injection defense). The pitch is that you can let agents touch internal systems with auditing and cost limits instead of hoping nothing goes wrong. Self-hosting is free for teams under 30 people. This is enterprise infrastructure, and it installs like it. Docker, Helm, and Kubernetes are the deployment paths, so standing it up is a platform-team job, not an afternoon. The upside of self-hosting is the whole point of the product: your prompts, your data, and your agent traffic stay inside your own boundary, which is exactly the property security teams want before they let an LLM near anything sensitive. Solo builders and small teams experimenting with agents can run it free, but it is heavier than you need unless governance is the actual problem you are solving. Where it earns its keep is the mid-size company standardizing how dozens of agents reach internal tools: the AGPL self-host covers you up to 30 users, and past that you are into enterprise licensing. The comparison set is commercial AI gateways like Portkey or Kong's AI Gateway; Archestra's bet is open source plus security as the differentiator. Two catches. It is young and venture-backed, which means fast movement but also a roadmap that answers to investors, so watch how the open-core line shifts over time. And the README's talk of migrating from Claude Cowork and similar reads more like marketing than the substance underneath, which is solid. Judge it on the gateway and guardrails, not the launch copy.
Kiro Crew keeps agent work running after you shut the laptop. A long-lived Gateway process holds sessions, memory, schedules and approvals, so multi-step tasks run unattended, recurring jobs fire on schedule, and heartbeats watch a system until a human is actually needed. Apache 2.0, on hardware you control. Getting it up is a one-line installer, a desktop package, a Docker image, or a source build wanting Python 3.12 and Node 22. You work from a dashboard on localhost, the desktop app, the CLI, or through Slack and Discord. History and memory stay on the host. The software is free. The models are not. Every path runs on kiro-cli, so model access goes through a Kiro account: $0 gets 50 credits a month on Claude Sonnet 4.5, Pro is $20 per user for 1,000, and the ladder climbs to $200 for 10,000 with overage at 4 cents a credit. The catch is in the project's own boundary document. Kiro Crew is deliberately KiroACP-only, the agent provider setting accepts exactly one value, and kiro-cli is required. An RFC asking to lift that is still in draft. The license buys the harness, not the choice of model underneath.
PilotDeck is an open source 'agent operating system' from OpenBMB, ModelBest, and Tsinghua's THUNLP, AGPL-3.0. It bundles three pieces most agent frameworks leave to you: a WorkSpace abstraction that keeps projects isolated, a white-box memory layer you can view and edit, and a smart router that sends cheap tasks to cheap models. Free, with a working web UI. Self-hosting is Docker, with TypeScript, Python, and Go components, so it's a real install, not a hobby script. The memory model is the most distinctive piece: instead of a black-box vector store you can't reason about, every entry is human-readable and auditable. The router claims around 70% cost reduction on real workloads by downgrading simple tasks to smaller models. For solo developers building agents who currently glue together LangGraph and a router, this collapses a few components into one and gives you a UI for memory and workspaces. Small teams running production agents get cost savings worth measuring. Larger orgs should treat it as research-grade until they verify the routing decisions on their own task mix. The catch: this is research-driven and the team is academic-plus-startup, so expect rapid changes and rough edges. The AGPL license also means anyone offering it as a service has to share modifications, which matters if you're embedding it inside a SaaS product. For internal use, it's an opinionated and ambitious starting point.
This gives you a local control center with full visibility. You get a dashboard that shows what OpenClaw is doing in real time, how much each task costs, and lets you set guardrails. It turns OpenClaw from 'fire and pray' into something you can actually trust and control. You see every API call, every decision branch, every token spent. You set budget limits, approve expensive operations, and kill tasks that go off the rails. MIT licensed, TypeScript. The catch: this is OpenClaw-specific. If you're using Claude Code, Cursor, or Codex, this does nothing for you. And 'control center' implies oversight, but you still need to understand what you're looking at. It surfaces the data, it doesn't interpret it for you. Early stage, so expect UI rough edges.
Mirage mounts S3 buckets, Google Drive, Slack, Gmail, and Redis side by side as one filesystem so an AI agent can use familiar Unix commands across all of them. Instead of teaching the agent five different SDKs, you point it at a virtual filesystem and let it grep, cat, cp, and pipe between services the way it would on a local disk. It's an abstraction layer designed for how agents already think.
Install is pip or npm or a curl one-liner. Python 3.12+ or Node 20+, macOS or Linux. You provide credentials for whichever backends you want to mount (AWS, Google, Slack, etc.) and Mirage exposes them as paths. There's no central service; everything runs locally inside the agent's environment.
Solo developers building agent workflows: this is the kind of glue you'd otherwise hand-roll, and having it as an Apache 2.0 package is useful. Small teams shipping agents: worth testing as part of your tooling stack. Large teams: monitor the project; it's young but the design is right.
The catch: v0.0.1, released May 6th, 2026. First public release. The abstraction is interesting but the implementation is brand new. Expect rough edges and breaking changes. Pin the version and read release notes carefully.
AgentENV runs thousands of isolated agent environments at once, which is exactly what reinforcement-learning training for agents needs. Written in Rust, it orchestrates Firecracker microVMs across a cluster with sub-50-millisecond snapshot pause and resume, memory forking, and S3-compatible storage. It comes from kvcache-ai, the group behind the well-regarded KTransformers, so the pedigree is real. The standout is the speed of the snapshotting. Forking agent state and pausing or resuming microVMs in under 50ms is what makes large-scale agentic RL practical instead of theoretical. This is free under MIT, but the audience is narrow: teams actually training agents at scale. Solo tinkering is possible, self-hosting on real infrastructure is the expectation. The catch is two-fold. It's infra-heavy, it wants Linux 6.8+ and access to /dev/kvm, so this is a datacenter or beefy-server tool, not a laptop one. And the README is blunt that the API has no authorization built in and must never be exposed publicly. Powerful and specialized, with sharp edges you have to respect.
Qclaw is a GUI wrapper for OpenClaw that removes the command-line barrier. If you want to use AI coding tools but the terminal feels intimidating, Qclaw puts a graphical interface on top of OpenClaw's capabilities. Chinese-language interface, built for users who prefer visual interaction over command-line workflows. It translates OpenClaw's CLI operations into clickable buttons and forms. The catch: Chinese-language only, and it wraps another tool rather than providing standalone functionality. You still need OpenClaw installed underneath. If you're comfortable with a terminal, OpenClaw directly is more flexible. And because it depends on another project's API, breaking changes upstream can break Qclaw.
Cindy is an open source AI agent that runs on your own machine and actually does the work, driving Claude Code or Codex, controlling a browser or the computer, and reaching into third-party apps. It ships as a desktop and mobile app with native binaries bundled in, so it's closer to a finished product than a framework you assemble. The client is Apache 2.0 and works fully local, including a "Skip Sign-In" path, so you can run it without an account and keep everything on your device. That's the free core, and for a lot of people it's the whole thing. Solo users and small teams can run it free against their own model keys. The monetization is an optional official Cindy service that bills model usage transparently, which is the sensible route for teams that want a managed backend instead of wiring up their own. The catch is maturity. Cindy is early and it shows, a large open-issue count and the rough edges you'd expect from a fast-moving agent app. The direction is good and the local-first stance is the right one, but test it on real tasks before you trust it with anything that matters.
Boxlite gives you lightweight sandboxes. Each sandbox is a stateful micro-VM with hardware isolation, snapshots, and an API to control it. Picture giving every AI agent its own disposable computer. The project is open source under Apache 2.0 and self-hosting is free. It's early but growing fast. The 'agent sandboxing' space is heating up as AI agents get more autonomous and need safer execution environments. The catch: this is emerging technology. The documentation and ecosystem are still maturing. Running Firecracker-based micro-VMs requires Linux with KVM support. No macOS, no Windows natively. And the question of whether you need full VM isolation versus Docker containers depends on your threat model. For most use cases, Docker is simpler. Boxlite is for when you can't trust the code being executed.
Codex-console is an integrated control panel for that workflow. Task management, batch processing, data export, auto-upload, log viewing, and packaging in one place. Built in Python with MIT license. The project provides compatibility fixes and experience optimizations for managing multiple concurrent AI coding sessions. The catch: the README and documentation are entirely in Chinese. If you don't read Chinese, you'll be navigating the tool through translation or code reading. The project is a console/dashboard wrapper. It doesn't do the AI work itself, it just helps you manage it. And at with limited English documentation, community support outside Chinese-speaking developers will be thin.
GitLab MCP fills the gap Anthropic left open. GitHub has an official MCP server for AI coding assistants. GitLab does not. This community-built server connects Claude Code, Cursor, Copilot, VS Code, and Codex to your GitLab instance, exposing merge requests, issues, pipelines, wiki, releases, and labels as callable tools. Setup is simple for local use: one npx command plus a personal access token. Self-hosted GitLab works fine with a custom API URL. For team deployments, there's a Docker image with OAuth2 support and multi-user remote authorization. Four auth methods cover everything from quick local testing to production multi-tenant setups. Solo developers on GitLab get AI coding assistant integration that was previously GitHub-only. Teams running self-hosted GitLab get the same MCP capabilities without migrating to GitHub. There's a read-only mode toggle for safety if you want to prevent the AI from making changes. The catch: community-maintained, not official GitLab or Anthropic. Feature parity depends on one maintainer keeping up with GitLab's API surface. The multi-user OAuth setup requires a public HTTPS endpoint and pre-registered GitLab app, which is non-trivial.
LLM Space is a desktop workbench for people building AI agents. Write the prompt, wire up the tools, run it, and watch every model call and tool invocation as it happens. Replay a failed run out of history and step through it. Threads are plain files on your machine and API keys stay local. MIT licensed and free, and the DeerFlow team says every release of bytedance/deer-flow is built and debugged inside it. There is no server to stand up. Download the DMG and go. Packaging is the constraint: the desktop app is macOS only, in a 27 MB edition on the system WebView or a 130 MB one with its own renderer. Linux gets a headless server build with no UI. It runs its own agent runtime, so it hosts your agent rather than observing one you already deployed. Use it while you are still shaping an agent and want to see what the model actually did instead of inferring it from logs. Solo builders: free, and the replay view alone earns the download. Small teams: free, though file-based threads make sharing a run manual. If you need hosted traces and production monitoring the whole team can see, langfuse/langfuse or LangSmith is the right shape. promptfoo/promptfoo fits better when evaluation is the main job. The catch: they only merge pull requests from the DeerFlow core team, so this is open source you can read and fork but not contribute to. Anonymous telemetry is on by default with a documented opt-out. And the public repo dates to mid-2026 even though the project claims a 2023 start, so treat the visible history as short.
This MCP server wraps 41 Brazilian public APIs into one standardized interface your agent can query. The smart part: it doesn't dump all 200+ tools on your agent at once. BM25 search filters to show only relevant tools per query, and a query planner can combine multiple APIs in a single call. 24 of the APIs need no authentication at all. The remaining ones use 2 optional API keys you get with free registration. MIT licensed. Built in Python with async httpx, Pydantic v2, and rate limiting with backoff. The catch: this is Brazil-specific. If you're not working with Brazilian data, there's nothing here for you. And wrapping government APIs means you inherit their reliability issues: downtime, rate limits, and data quality are the API provider's problem, not Floci's. The project is brand new and maintained by what appears to be a single developer.
OpenSquirrel is a native desktop app that puts Codex, Cursor, and OpenCode in one window so you stop losing track of what each one is doing. A control plane for your AI coding agents, built in Rust with the GPUI framework. What's free: Everything. MIT licensed, fully open source. No paid tier, no cloud service. The pitch is honest: you're squirrely, you jump between agents, and you need a way to see them all at once without alt-tabbing through six terminal windows. The Rust/GPUI foundation means it's fast and native, not an Electron wrapper eating 2GB of RAM. The catch: this is early, so it's not battle-tested yet. GPUI (Zed's UI framework) is relatively new itself, so you're building on new foundations. If you only use one AI coding tool, this adds zero value. It's specifically for the multi-agent workflow that a growing number of developers are adopting.
Commonly gives coding agents a shared workspace instead of a fresh amnesiac session every time. Claude Code, Cursor, Codex, and your own agents each get persistent identity, their own memory, their own skills, and a workstation, and you talk to them in a chat room where the work actually accumulates. The problem it targets is real and annoying: every new agent session starts by re-explaining the project. Commonly keeps that context attached to the agent rather than to a terminal tab you closed. Task management and cross-runtime handoff sit on top. Installation is Docker with the Compose v2 plugin, then clone and run the install script, and you are on localhost:3000. Apache 2.0, self-hosted by default, with no per agent fee and no seat pricing. There is a live demo at commonly.me if you want to see it before committing a machine to it. The catch is that the README says the project is early, and it means it. This is a category where the shape of the right answer is still being argued about, and multi-agent workspaces have a habit of being rewritten. Self-host it, keep your data local, and do not build a team process on it yet.
VM0 runs AI coding agents in isolated cloud sandboxes on a schedule. Describe a workflow in natural language, point it at a repo, and it executes in a Firecracker microVM with full Claude Code compatibility. Think of it as cron for AI agents, with sandboxing built in. The platform gives you persistence (resume, fork, version sessions), observability (logs, metrics, network visibility), and integration with 35,000+ skills via the skills.sh ecosystem. Self-hosting means running the Firecracker VM infrastructure yourself, which is a real infrastructure commitment. Solo developers who want to automate repetitive coding tasks (daily CI fixes, dependency updates, report generation) get the most value here. Teams running multiple agents benefit from the orchestration layer. The catch: this is very early stage. The license isn't a standard OSS license, the docs are sparse, and you're building on a startup's roadmap. The managed cloud is the realistic path for most users, and pricing for that isn't finalized yet.