7 open source tools compared. Sorted by stars. Scroll down for our analysis.
See our ranked picks: Best Open Source Claude Code & Codex Skills
By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.
| Tool | Stars | Velocity | Score |
|---|---|---|---|
codex-plugin-cc Use Codex from Claude Code to review code or delegate tasks. | 33.6k | +398/wk | 78 |
jev-ultrafast i. am. speed. | 21.0k | +4196/wk | 88 |
OpenBot Open-source AI coworkers that each get a computer of their own: a browser, files and tools, with every action decided before it happens and recorded after. Bring any AG-UI agent. | 5.7k | +367/wk | 82 |
skill-recorder Desktop app that records your on-screen work session and uses the GitHub Copilot CLI to reconstruct it as an intent + ordered steps, then builds a reusable Skill or Automation for Microsoft Scout, Microsoft Copilot Cowork, or Copilot Studio. | 4.1k | +168/wk | 78 |
chrome-cdp-skill Give your AI agent access to your live Chrome session. Works out of the box, connects to tabs you already have open | 3.3k | +6/wk | 62 |
phone-harness let your agent control your phone | 3.1k | +110/wk | 72 |
chromex A Codex-powered Chrome side-panel assistant for page context, tabs, voice, and image workflows. | 1.2k | +4/wk | 55 |
Stay ahead of the category
New tools and momentum shifts, every Wednesday.
OpenAI's Codex agent packaged as a Claude Code skill plugin. It lets you invoke Codex from inside Claude Code to review code or delegate tasks, connecting two AI coding agents so they can collaborate. Useful if you want a second opinion from a different model without switching tools. The integration is straightforward: install the skill, and you can ask Claude Code to hand off specific tasks to Codex. Code review is the primary use case, where having two different models look at the same code catches more issues than either alone. The catch: requires both Claude Code and OpenAI API access, so you're paying for two AI services to talk to each other. The value proposition only makes sense if you're already invested in both ecosystems. For most developers, one AI coding tool is enough.
jev-ultrafast is Browser Use's speed demo for a new kind of model. Normal browser agents ask a chat model to write out its next step. This one hands Jev, a decision model from TypeSafe AI that returns typed choices instead of text, a numbered list of page elements and gets back which one to click. A small text model only writes the words that need typing. The README's showcase is a Zurich to London search on Google Flights in about seven seconds. MIT licensed. It is source code, not a product: uv sync, a local Chrome with remote debugging on, and two paid keys. One is a TypeSafe key for Jev, which bills $0.042 per million input tokens with free output and is waitlisted right now. The other is an OpenRouter key for the text model. No releases, version 0.1.0, and the README calls it an MVP. Solo builders curious about decision models: worth an evening once you clear the waitlist. Teams that need browser automation today should use the same company's main browser-use library, or Skyvern. The catch: the benchmark is three repeats of one task on one browser profile, and the README says so. Shadow DOM, iframes, canvas, uploads, and pop-up tabs are out of scope. And the whole thing depends on a closed API you cannot self-host.
OpenBot gives every agent its own computer. Not a sandboxed shell, an actual browser with its own logins, its own files, and only the tools you granted it, so the agent that books travel never holds the credentials that could send email. Every action is decided before it runs and recorded after, which is the part most agent frameworks skip. MIT licensed, from the CopilotKit team. Any AG-UI agent plugs in, written on a framework or by hand, and shows up as a coworker with its own channel. You watch it work on its own screen, take the wheel when it reaches something it should not do alone, and hand control back. It answers with rendered components rather than only prose, which makes reviewing its work faster than reading a transcript. Running it means containers, browser sessions, and credential storage per agent. That is real infrastructure, not a CLI. Solo builders can run it locally. Teams should plan for isolation between agents from the start. The catch is age. This shipped in August 2026 and moves daily. The permission model is the whole value proposition here, and a permission model with a week of production history has not been tested by anyone who wanted to break it.
Skill-recorder watches you do a task once and turns it into a skill your AI agent can repeat. It's a Microsoft desktop app that records your screen session (window switches, URLs, optional spoken narration), then uses the GitHub Copilot CLI to reconstruct what you did as an intent plus ordered steps, packaged as a reusable skill file or a scheduled automation. MIT licensed and free. The smart design choice is that it generalizes instead of replaying. Show it one form submission and it writes a procedure for handling forms, reaching for native tools like the GitHub CLI instead of scripting browser clicks. Recording and Whisper transcription stay on your machine, but hitting Analyze ships window titles, URLs, clipboard previews, and screenshots to GitHub's cloud. The README tells you flat out not to record credentials. That warning is honest, and it also rules out most real admin workflows. This only makes sense inside Microsoft's agent ecosystem. The output targets Microsoft Scout, Copilot Cowork, and Copilot Studio, and you need GitHub Copilot access to run the analysis. If you live in that stack, install it. If you don't, the skills it produces have nowhere to go. The catch: the recorder is free, the destinations aren't. Copilot Cowork needs a Microsoft 365 Copilot license at $30/user/mo plus usage billing, and the app is weeks old. MIT on the recorder doesn't buy portability of the result.
This skill connects your agent to your live Chrome via the Chrome DevTools Protocol (CDP). Your agent can read pages, click buttons, fill forms, and navigate, in the browser you're already using. The difference from tools like Playwright is that this connects to existing tabs. Your agent can interact with pages where you're already authenticated, see what you see, and do what you'd do manually. MIT licensed, JavaScript. The catch: giving an AI agent access to your live browser session with all your logged-in accounts is a real security consideration. The agent can see everything you can see, including sensitive data in open tabs. There's no permission model beyond 'full access.' And CDP connections can be fragile; Chrome updates can break the protocol.
Phone-harness lets an AI agent physically drive your iPhone. It captures the macOS iPhone Mirroring window, runs Apple's OCR over it for eyes, and posts real taps, drags, and keystrokes for hands. No jailbreak, no Xcode, about 500 lines of Python, MIT licensed and free. Setup is a one-time iPhone Mirroring pairing plus Accessibility and Screen Recording permissions, and a doctor command checks the chain. The design constraint is that the mirroring window is just a video stream: no accessibility tree, so anything not rendered as readable text is invisible to the agent, and the window must stay frontmost, so your Mac is occupied while it runs. This is a proof of concept for people experimenting with agent-driven mobile automation, not a QA tool; Appium and Maestro remain the serious answers for testing. And it does not work in the EU at all, because Apple has never shipped iPhone Mirroring there. The catch: you're giving an LLM unsupervised control of your personal, logged-in phone with nothing sandboxing it. Very young, tiny commit history, and the issue list is already outrunning the code. Fun to try; think hard before trusting it.
Chromex is a Chrome side-panel extension that connects your browser to OpenAI's Codex through a local native messaging bridge. Summarize pages, work across tabs and screenshots, edit images, transcribe voice, and run browser-control workflows with visible in-page indicators. MIT licensed.
Setup is heavier than a typical extension: clone the repo, npm install && npm run build, run install-native-host.mjs, then load the unpacked extension at chrome://extensions. The architecture (Chrome extension to native host to local bridge to codex app-server) keeps your API key local; raw keys aren't stored in extension storage.
Pick this if you live in Chrome, already run Codex, and want one assistant that sees the page you're on. Solo: free, you pay only for Codex tokens. Small teams: same. Large teams or non-Codex shops: skip; this is built around Codex specifically.
The catch: Codex-only. Switch to Claude or Gemini for your CLI agent and Chromex doesn't follow. The native bridge is only as polished as the project, which is small and early. For a more mature Chrome AI assistant, Sider and Monica have more features and broader model support.