How to Build an Agentic OS (Claude OS) That Actually Works
How to Build an Agentic OS (Claude OS) That Actually Works
You need an Agentic OS — but not for the reasons you think. It's not the dashboards or the pretty visual interfaces. It's about building a custom system that takes the best of every AI tool you have — GPT-6, Astra, Claude Opus 5.5, Jev — and builds it around you and your actual work in a way you can't get inside the terminal. I've built exactly this for my own business, and in this post I'll show you how it works and how to build one yourself.
An Agentic OS (or Claude OS) is a three-layer system: a visual layer, a memory layer, and a skill layer. Most people obsess over the first one. The real value lives in the other two. Here's how all three fit together — and why the skill backbone is the one thing you can't skip.
What Is an Agentic OS?
An Agentic OS is a custom system that wraps around the AI tools you already use and organizes them around how you actually work. It has three layers:
- Visual layer — the dashboard you see and interact with. Metrics, calendar, skills you can run with a button, research tabs. 100% custom to you.
- Memory layer — an organized file-and-folder structure (I use Obsidian) that lets your AI find answers fast and accurately across thousands of files.
- Skill layer — the backbone. Every repeatable task in your life or business, codified into skills and automations that produce deliverables.
The key idea: the deliverables and outputs your skills generate get fed back into the OS. All the reports, all the graphics, all the metrics on the dashboard are driven by the skills you run. It's a loop, not a pile of disconnected tools.
And there are no rules about what the visual layer has to be. Mine runs as a web app in the browser, but I've also ported it into Obsidian as a plugin. Both can run on Claude Code or Codex, both give me a terminal, both let me click a button to fire off any skill. The point is it bends to your workflow, not the other way around.
How Does an Agentic OS Route Tasks to Different Models?
This is where it gets interesting. When you talk to the OS, it routes your request to a different model depending on how complex the task is. I use Jev as the router — a classifier that sorts every request into one of three tiers.
Tier 1 — instant, no AI needed. Simple commands like "bring up the morning intel brief" or "open the terminal." Jev classifies these and the system executes them directly. No LLM is even called, which is why the response is near-instant — we're talking hundreds of a millisecond.
Tier 2 — light thinking. Questions like "what was the biggest news in AI today?" Jev routes these to the smallest available model — Haiku or Luna, depending on whether I'm running the Astra version or the Claude Code version. Cheap, fast, and if you wanted a fully local setup you could swap in a model running on your own machine.
Tier 3 — real work. "Create a visual explainer illustrating the difference between Jev and standard LLMs like Fable." This pulls up Claude Code or Codex in your terminal and does the actual heavy lifting.
The whole flow, end to end: I speak a request. It goes to my Obsidian plugin (or the Jarvis-style HUD). That goes to a bridge, which records everything in Obsidian and engages the local voice system. Whisper — open source — transcribes my speech into text, acting as the model's ears. The transcription gets sent to Jev, which classifies the tier. The voice you hear back is Kokoro, another open-source model you can swap for any voice you want. Every conversation and deliverable gets logged back into Obsidian.
All the speech, transcription, and audio happens locally on your machine. That's the part most "voice agent" setups get wrong — they ship your audio to some API. This doesn't.
What Is Jev and Why Use It for Model Routing?
Jev lets you ask an AI model fuzzy questions and get specific numbers back — probabilities, not sentences. It's an AI system, but it's not an LLM, which is why the discourse around it is so confusing. It's not a replacement for Claude Code or Codex in general, even though people frame it that way.
Here's the example that makes it click. Imagine a customer service scenario. A customer writes: "I canceled my subscription last month but you charged me again." You need to know which team should handle it.
An LLM handles this fine — it'll say "this is a billing issue, route it to billing." Correct answer. Jev also gets the correct answer, but instead of a sentence it returns a probability: 92% chance this should go to billing. No back-and-forth, no chat.
So if both get it right, why care? Because Jev does it roughly 200x faster and cheaper. For any "what should I do here, given a fixed set of choices" problem — exactly what model routing is — Jev is perfect. The predefined choices in my router are the three tiers. I could put Haiku or Luna in charge of that classification, but there's no point anymore. That's why Jev is integrated.
Why Is the Skill Backbone the Most Important Layer?
If you ignore everything else in this post, do this one thing and you'll be ahead of 99% of people using Claude Code or Codex. Take every task you do day-to-day and week-to-week — personal or business — and turn it into a skill. If it makes sense, turn that skill into an automation.
The value seems obvious, yet almost nobody does it systematically. They codify a couple of tasks and stop. The reason is it feels like too much work and you don't know where to start. Good news: it's actually easy, and there are only two ways in.
Method 1 — mine your logs. Claude Code and Codex keep a written history of every conversation you've had for the last 30+ days. Instead of guessing how AI should help you, point it at reality: "Based on everything I've done over the last 30 days, are there any skills we could create?" That's the entire prompt. No special workflow.
Method 2 — the brain dump. Sit in front of Claude Code or Codex, open your mic, and talk non-stop for 10 to 20 minutes about what you do day-to-day and week-to-week. It doesn't have to be coherent — stream of consciousness is fine, the model works through it. Then ask: "Based on everything I just told you, what skills or automations could take work off my plate?"
Neither method requires you to have the answers walking in. You don't need to know what should become a skill — that's the point. Either approach gets you 90% of the way to a set of skills that make your life dramatically easier.
Deciding which skills become automations is usually straightforward, and you just ask Claude to wire it up. You can do it manually through routines or scheduled tasks in Codex, or let the system trigger them off your machine directly. Lots of flexibility, and again — you don't need the answer before you walk in. You just have to ask.
How Does Obsidian Work as the Memory Layer?
Everything that goes into or comes out of the OS gets recorded in the memory layer. I use Obsidian, a free, simple desktop app for navigating markdown files. Paired with Claude Code or Codex, this is what people call a "second brain."
But here's what most people get wrong: Obsidian doesn't give your AI any magic memory upgrade. It's not a RAG system. The knowledge graphs look cool but that's about where the functionality ends. Using Obsidian doesn't make Claude remember better.
What Obsidian actually does is make it easy to navigate the files and folders you've designated as your vault. That's a subtle but critical distinction: the power of a second brain is 100% in how you organize your files and folders, not in Obsidian itself. Dump 10 million random files into one unorganized folder and Obsidian helps you exactly zero.
So you need a structure that makes sense. Mine follows the Karpathy setup — Andrej Karpathy coined this pattern for Obsidian + Claude Code. It's simple. One vault folder, three subfolders underneath:
- raw — raw data goes here
- wiki — where you organize raw data into Wikipedia-style articles
- outputs — finished deliverables like slide decks
A concrete flow: I ask the OS to research AI agents. It downloads a pile of information into raw. I tell it to organize that, and it creates an AI-agents subfolder in wiki with clean articles — autonomous coding, tool-use patterns, and so on. Then I say "make a PowerPoint about this," and the deck lands in outputs.
Because it's organized this way, when I ask Claude about anything in my vault it has a clear map and a clear path to find it — fast and efficient, no crawling through ten million folders. And it's just as easy for me, the human, to find things.
Does it have to be raw/wiki/outputs? No. Set it up any way that makes sense as long as it's navigable by both you and your AI. And you don't need the answer up front — tell your OS "this is my vault folder, help me structure it" and it'll map it out and move files as needed. The one thing I strongly recommend: put a CLAUDE.md or AGENTS.md in your vault that documents the structure and the rules for adding new files — so the system never drifts out of whack over time.
Why Does the Visual Layer Matter If You Can Use the Terminal?
Every skill produces deliverables. Instead of scattering them across a dozen places, the visual layer consolidates them into one customized dashboard — metrics, schedule, tasks, morning headlines, audience numbers, deeper research. That alone is hard to do cleanly inside a desktop app and a genuine pain inside the terminal.
But the visual layer's real payoff shows up when you bring in team members and clients. If you're technical enough to mash something together in the terminal for yourself, fine — but you can't easily hand that to a client. A custom dashboard you can.
Here's the mechanism: every button on the dashboard, when clicked, fires a headless version of Claude Code or Codex in the background to execute that skill. So I can sit someone in front of my dashboard — the Obsidian plugin or the web app — and have them run skills that do serious work, harnessing 90-95% of the power of Claude Code or Codex without ever knowing how to open it.
That's why this isn't a one-size-fits-all product. Whatever you need in one place, you build it that way. And you can package it — for teammates, or to sell to clients.
How Do You Start Building Your Own Agentic OS?
Start from the ground up with the skill backbone — it's the highest-leverage piece and the only one you truly can't skip. Use the log-mining prompt or the 10-20 minute brain dump to generate your first batch of skills, then ask Claude which ones should become automations.
Next, set up the memory layer. Point your AI at a folder, tell it to help you build an organized structure, and add a CLAUDE.md documenting the rules. Then, if you want it, build the visual layer to consolidate the deliverables your skills produce.
You don't need to have the answers before you begin. The recurring theme across all three layers is the same: you just have to ask Claude or Codex to do it.
Frequently Asked Questions
What is an Agentic OS?
An Agentic OS (also called a Claude OS) is a custom system with three layers — visual, memory, and skills — that wraps around your AI tools and organizes them around how you actually work. It routes tasks to the right model, stores everything in an organized memory layer, and surfaces the results in a custom dashboard.
Is Jev a replacement for Claude Code or Codex?
No. Jev is a classifier, not an LLM. It returns probabilities instead of sentences and is roughly 200x faster and cheaper for "which of these fixed choices applies" problems — like routing a request to the right tier or model. It complements Claude Code and Codex; it doesn't replace them.
Does using Obsidian make Claude's memory better?
No. Obsidian is not a RAG system and gives the AI no memory upgrade. It just makes an organized file-and-folder structure easy to navigate for both you and your AI. All the power comes from how you organize the vault — not from Obsidian itself.
How do I decide what to turn into a skill?
You don't have to decide up front. Either point Claude Code or Codex at your last 30 days of conversation history and ask what could become skills, or talk for 10-20 minutes about your day-to-day work and ask the same question. Both surface the right candidates without you knowing them in advance.
Do I need a dashboard to have an Agentic OS?
No. The visual layer is optional and mostly pays off when you're bringing team members or clients into the system. The skill backbone is the non-negotiable part — the memory and visual layers build on top of it.
If you want to go deeper into building your own Agentic OS, join the free Chase AI community for templates, prompts, and live breakdowns. And if you're serious about building with AI, check out the paid community, Chase AI+, for hands-on guidance on how to make money with AI.


