
Alright, so you want to build applications with AI. Not just little toy scripts, but real, production-grade software that can actually improve itself over time. That’s the whole premise behind what I call a software factory, and if you’re serious about AI-driven development, you need to understand how this works.
The core idea is mimicking a real software development team but replacing the humans with specialized AI agents. You’ve got a builder, a QA specialist, a reviewer, and the whole cycle is orchestrated through GitHub issues. A ticket gets created, the builder agent works on it, QA verifies it, and if it doesn’t pass, it gets kicked back for revisions. Every single action, every pull request, every comment from an agent gets logged right there in the GitHub issue. So a human can drop in at any point, look at the issue log, and see exactly what the AI did and why.
I’ve been using this agentic framework for months to build applications, fix bugs, and even automatically reduce user churn. There are guard rails and hooks in place to make sure the agents are building things the right way. And I’ve open-sourced the entire thing: the software factory, the skills, everything. Each skill is either one I’ve built and used in my day-to-day work or one I’ve carefully reviewed on this channel. You can even run the factory on any model you like.
How the Cycle Actually Works
The whole process kicks off with telemetry. We’re talking tools like Sentry or PostHog that monitor your system performance and user analytics. A skill called Super Collect runs on a schedule, grabs those issues, and adds them as tickets to your GitHub project board.
From there, another skill called Super Board takes over. This is where the factory of specialist agents lives. Tickets get processed one by one on the kanban board. If a ticket needs human approval, it simply gets assigned to you and blocks there until you sign off.
One thing I want to emphasize: the models are easy to swap. Whether you’re using Claude Code, Codex, or something else entirely, you can switch out the underlying model without rearchitecting the whole system. The setup is designed to be dead simple. You run the onboarding, connect your GitHub account, point it at the right project board and branch, set your system prompts and policies, and you’re good to go.
Inside the Specialized AI Agents
The big-picture flow is one thing, but the real power is in the specialized agents themselves. I’ve reviewed tons of spec-driven development skills on this channel, and I’ve curated the best ones into these agents so they actually handle things properly.
When you run Super Board, it triggers the build process the moment a ticket lands in the builder column. But before a single line of code gets written, two preparation skills fire first.
Preparation: PonyTail and Context7
The first skill is called PonyTail. This one is incredibly popular, over 100,000 stars on GitHub. Its job is to reduce code complexity in your codebase because, believe it or not, a lot of AI agents will write code that’s just duplication after duplication. A single file can balloon to a thousand or even a hundred thousand lines.
PonyTail uses a ladder of seven checks to decide if code should even be written at all:
- Does this feature need to exist?
- Does something in the codebase already do this? Maybe a utility function or a component we can reuse.
- Is there a standard library that solves this?
- Is there a native platform feature we can use right away?
- Is there an already-installed dependency that handles this?
- Can it actually be done in one line?
This drastically reduces code duplication and also the tokens the model consumes. The more code it writes, the more tokens it burns and the more time it costs. If an existing solution exists, use it.
After PonyTail, we have Context7. Chances are your application uses a bunch of dependencies: Vercel, React, NestJS, whatever third-party libraries you’re pulling in. The AI needs to know their actual documentation before it writes a line of code. Context7 reads through those docs so the agent understands exactly how to use the libraries correctly.
Building with Map Pock Skills
Once the prep work is done, we move on to the actual building process. This is where Map Pock skills come in. I’ve reviewed so many spec-driven development skills: GSD, GSTA, Stack Superpowers, and others. Map Pock is lightweight and really follows best practices for utilizing the model’s capabilities.
It handles three distinct scenarios:
- Implements for new features
- Diagnosing for bug fixes
- Codebase designs for refactorings
I’ve made a full tutorial breaking down all the Map Pock skills if you want to go deeper.
Checks and the Humanizer
After the building phase, we run checks: verifications, code reviews, and then a skill called Humanizer. This one makes sure the AI doesn’t write overly complex words that make the code harder to understand. Once that’s done, the handoff phase wraps everything up and posts it back to the GitHub issue.
Super QA: Testing the Right Way
This is the most important part. Super QA focuses on verifying actual ticket completion. As someone who worked as a senior software engineer at companies like Amazon and Microsoft, I can tell you that testing is critical. A small feature, once shipped, can be used by millions of people. One bug can cause real financial damage in a production application.
What makes this QA skill different from other testing skills is that it uses the proper testing libraries in the right order. Most people jump straight to end-to-end testing with Playwright, but that’s the last step. You need to start at the bottom layer first: fundamental unit testing. Then move up to integration testing or component testing. Test your actual React components because they depend on each other. Components call different functions and hooks, so you need to make sure the bottom layer is fully solid before you move up.
Only once component testing is done do you move to Playwright or Cypress for end-to-end UI testing. For backends, there’s a whole other set: stress tests, integration tests, and more. I’ve embedded this entire knowledge hierarchy into the Super QA skill so it knows exactly which testing library to use for which fix.
Super Review and the Refinement Loop
The review process brings PonyTail back into play for a review pass, along with a full code review and codebase design check to make sure everything is properly refactored.
Then there’s Super Collect, which gathers tickets from different sources like Sentry or PostHog. But it doesn’t stop at bugs. If you want to brainstorm new features, you can use skills like Wayfinder or grooming skills to brain-dump ideas into GitHub issues. Those ideas go through the same verification process and land in your backlog.
On top of all that, I’ve built a custom skill called UI Refinement Loops. Let’s say you have a page or component that looks like AI slop. How do you get it to self-improve in a loop? This skill depends heavily on Map Pock skills and another one called Impeccable. Impeccable is specifically designed to help create UI designs that are less AI slop.
Here’s how it works: you give it a direction, like a page or component redesign, and point out the specific problems you see. It asks you a couple of clarifying questions about the direction you want to go, then runs in a loop. First it diagnoses, critiques, and audits. Then it moves through fixing, hardening, and clarifying. You can read more about the exact flow in the documentation, but it also handles styling, adding motion, polishing, and then circles back to diagnose again to see if there are more problems to fix. It usually runs a few iterations, and you end up with a much better UI design than what you started with.
There are also sub-skills like Git Sync that resolve conflicts, pull changes, and write proper commit messages using best practices. And guard rails are in place to block secrets from being written to repositories, automatically set up skills eval files, and update any outdated documentation.
Points clés à retenir
- A software factory mimics a real dev team with specialized AI agents for building, QA, and review, all tracked through GitHub issues.
- PonyTail reduces code duplication before any code is written by checking if a solution already exists in the codebase, standard libraries, or dependencies.
- Context7 reads third-party documentation so the AI knows exactly how to use libraries before writing code.
- Map Pock skills handle feature implementation, bug diagnosis, and refactoring in a lightweight, best-practice way.
- Super QA uses a layered testing approach: unit tests first, then component tests, then end-to-end tests, with the right library for each layer.
- UI Refinement Loops run multiple iterations of diagnose, critique, fix, and polish to turn AI slop into solid UI designs.
- The entire system is open-sourced, model-agnostic, and designed to be set up in three steps.
I’ve packaged seven years of my domain knowledge as a senior software engineer into these skills, along with all the curated skills I’ve reviewed on this channel. The setup process is kept really simple: install it, run the onboarding so every dependency is connected, then just run Super Board. It fans out multiple agents, drains the tickets in the ready column, and gets to work.
If you want to see the full lifecycle of how teams at Anthropic actually build applications, I’ve made a video on that too. But for now, that’s the software factory.
