Where software is built by a governed team of AI agents — and the framework they follow is the product.
I’m Marco. I build systems where AI agents develop software through formal pipelines, quality gates, and specifications that live in git. Not prompts and prayers. Governed engineering.
development
AI agents
processed
tracked
the codebase
without a spec
Three tracks, one principle: everything as code.
Agentic Flow Framework
A governed multi-agent SDLC pipeline. 12 specialized AI agents develop software through a four-layer process — Business, Architecture, Implementation, Verification — with human-approved quality gates between each layer.
The framework manages its own evolution through the same pipeline it uses for application code. The specifications, the governance rules, the agent profiles — they all live in git, versioned and immutable.
Everything as code. Not just infrastructure. Everything.
Empirical SDLC Research
I’m building SDLC-bench — the first benchmark that measures process quality in governed agentic development, not just whether an agent can fix a bug.
Seven scoring dimensions. Ablation studies that quantify which governance components actually matter. The goal: an academic paper and a benchmark the community can use.
Because you can’t improve what you can’t measure, and nobody is measuring this yet.
Weekly Research Blog
Deep dives on governed AI development. What works, what doesn’t, what the data says. No marketing fluff, no “10x your productivity” promises. Just the work.
Written by a human-agent team: I set the direction, Claude does the heavy lifting, and we argue about architecture decisions in commit messages.
A team of one human and twelve agents.
macrocode.ai is not a consultancy. It’s not an agency. It’s a workshop.
I’m a solo founder who builds with AI agents the way a recording artist works with a band — I write the songs, set the tempo, and make the calls. The agents play their parts. Each one is a specialist. None of them freelance. They follow a pipeline. They read specifications before they write code. They compile and test before they return. When they break a rule, there’s a violation report.
Claude is my co-author, my pair programmer, and occasionally the one who tells me my commit messages are too terse. In the blog and the dev log, you’ll read both our voices. When I say “we,” I mean it.
Marco Mancuso
Solution architect by background — microservices, DevOps, Agile (Scrum and SAFe certified), deep in the Atlassian ecosystem for years. I’ve spent my career in the space where complex systems meet disciplined engineering.
macrocode.ai is where I build what I can’t find: governed development tooling for the AI era. I started it in 2020, and the thing I’m most proud of is that every claim on this page comes from a git log, not a pitch deck.
I’m an introvert. For years I built in silence — the natural state for someone who’d rather read a paper on distributed consensus than pitch a slide deck. But what I’m seeing in agentic development is too significant to keep to myself. The industry is adopting AI coding tools at speed, and almost nobody is talking about governance, traceability, or process quality. So I’m sharing the journey — the data, the failures, the architecture decisions — because the only way to build a discipline is to build it in public.
If you’re an engineer who cares about how AI agents should be governed, a researcher studying agentic software development, or just someone who thinks the current “vibe coding” wave needs an engineering counterweight — I’d like to hear from you.
When I’m not building frameworks, I play guitar, lose myself in video games, read about physics, and stare at math until it becomes beautiful. I believe complex systems have an aesthetic — and that the best engineering is indistinguishable from art.
Building in public. Looking for fellow builders.
I’m documenting every step of this journey — the governed pipeline, the agent architecture, the research, the mistakes. Not because I have the answers, but because I think the questions matter and nobody else is asking them at this level of rigor.
If you’re working on governed AI development, studying agentic systems, or just tired of the “AI will replace developers” noise and want to talk about how AI agents should actually be engineered — let’s connect.
Blog – Latest Articles
Open-Sourcing the Agentic SDLC Framework: v2.0 Is Here
Open-Sourcing the Agentic SDLC Framework: v2.0 Is Here After 12 weeks of research, benchmarking, and community building — the Agentic Flow Framework goes open source. Here is everything we learned, what v2.0 includes, and how [...]
The Negative Self-Referential Tax: When Governance Overhead Becomes a Benefit
The Negative Self-Referential Tax: When Governance Overhead Becomes a Benefit Conventional wisdom says governance adds overhead. Our data shows the opposite — framework changes processed through the governed pipeline cost less than application changes. The [...]
From Single Project to Starter Kit: Extracting a Governed Framework
From Single Project to Starter Kit: Extracting a Governed Framework The hardest part of open-sourcing an internal framework is separating the generic from the specific. We document the extraction process — coupling audits, configuration layers, [...]
The Knowledge Architecture: Hierarchical Context Injection for Multi-Agent Systems
The Knowledge Architecture: Hierarchical Context Injection for Multi-Agent Systems When twelve AI agents write to the same knowledge base concurrently, every architectural shortcut becomes a failure mode. A flat directory of markdown files works for [...]
Mechanical Enforcement and the Irreducible Human Gate
Mechanical Enforcement and the Irreducible Human Gate You can tell an AI agent to follow the rules. You can put the rules in its system prompt, in its context window, in a file it reads [...]
The Linguistic API: Process-Theoretic Interfaces for Agent Governance
The Linguistic API: Process-Theoretic Interfaces for Agent Governance In every multi-agent system we have built or studied, the orchestrator receives its instructions as prose. Natural language rules in markdown files. "Follow the four-layer pipeline." "Check [...]
Gemini Deep Research Report: Macrocode.ai Analysis (Unaltered)
Gemini Deep Research Report: Macrocode.ai Analysis (Unaltered) This is the complete, unaltered output from Google Gemini Deep Research when asked to produce an “industry grade, multidimensional, thorough, structured and honest analysis” of macrocode.ai. Published verbatim [...]
The Immutability Illusion
The Immutability Illusion: Why Your Compliance Toolchain Cannot Guarantee What It Promises In our first post, we introduced the Agentic Flow Framework and its four-layer pipeline. We mentioned, almost in passing, that the framework stores [...]
Coming soon…
Agentlog – Latest Entries
DL-025: Progression Is a Graph, Not a List
DL-025: Progression Is a Graph, Not a List Most games store progression as a list — level 1, level 2, level 3. This one stores [...]
DL-024: The Editor Is the Compiler
DL-024: The Editor Is the Compiler A node graph you wire on a canvas, then press Run and watch the output render live inside the [...]
DL-023: The Same Algorithm Made Three Different Things
DL-023: The Same Algorithm Made Three Different Things A 2007 Eurographics paper on growing trees. A 1964 Japanese paper on water transport in plant stems. [...]
DL-021 Part 2: The Rule That Caught Itself
DL-021 Part 2: The Rule That Caught Itself Everything went wrong, all at once, and every single failure was the pipeline catching itself doing the [...]
DL-021 Part 1: The Content Engine
DL-021 Part 1: The Content Engine We set out to publish yesterday's devlog. The website caught a compliance gap, the wrong fix took the site [...]
DL-020: The Kit of Parts
DL-020: The Kit of Parts I wrote a publication structural integrity gate in the morning, an honest pipeline assessment in the afternoon, and a same-day [...]
DL-019: The Day We Paid the Drift
DL-019: The Day We Paid the Drift Three days of governance drift compounded into a single session of formalization. We meant to resume the Italian [...]
DL-018: The Five Lenses
DL-018: The Five Lenses Three documents in an inbox, two research ideas, five independent assessors, four new research entities. The framework learned to judge research [...]
DL-017: The Twenty-One Hour Day
Blog Post DL-017: The Twenty-One Hour Day Four initiatives in one session. A contamination incident, a translation pipeline, three fabricated legal claims, and a brand [...]
DL-016: The Compliance Session
Blog Post DL-016: The Compliance Session A legal obligation turns into a brand identity. The orchestrator builds a compliance pipeline, discovers the site collects nothing, [...]
DL-015: The Missing Layer
DL-015: The Missing Layer The orchestrator ships v1.0.0, gets corrected three times on the same concept, and discovers the framework has been flying without instruments [...]
DL-014: The Inbox
DL-014: The Inbox A forwarded tweet enters the pipeline. The pipeline processes it, fails at quality, and learns something about itself. The Morning Something arrived [...]


















