Agent Mesh

A multi-agent network where Claude Code and Gemini coordinate over SMTP and IRC. No custom API meshes. No complex coordination layers. Just battle-tested protocols, persistent semantic memory in Qdrant, and a subscription-based model with zero API token burn.

Where This Started

I started by trying to get my AI projects to talk to each other. What I ended up building is a simple multi-agent network: Claude Code in CLI, one for each project, Gemini in CLI handling QA and research, all writing to Qdrant for stateful memory. Subscription based. No API token burn.

The pattern AgentMesh is built on event-driven agents, local models for routine work, frontier models only for reasoning, persistent semantic memory that survives session boundaries — turned out not to be novel. Anthropic shipped two native API features implementing the same primitives the same week I finished building. I didn’t copy them; they shipped after my design was locked. I mention it not as validation, but because independent convergence on the same patterns suggests the patterns are probably right.

The institutional memory of the system stays in your infrastructure. By design you are not locked into one vendor’s model, different models, different blind spots, address different gaps.  Any LLM, Close one session, open another, memory persists. Simple is the point.

  • MD files = identity/role at boot (read fresh every session start)
  • Email = transport + intent (and gets logged to Qdrant)
  • Qdrant = the durable historical record (everything flows in)
  • IRC = in-flight only, ephemeral by design
  • The Brain = implements three orchestration primitives: dispatch (send work to a named worker), liveness check (verify the worker exists before dispatching), and fallback (route to an alternate worker if the primary cannot be reached). These are the same primitives a container orchestrator implements for containers. AgentMesh implements them for AI agents using shell scripts and tmux session names.

What I Built: The Five Layers

AgentMesh is five layers and each has a single job. Everything runs on two pieces of hardware I already owned: a Mac Mini M4 Pro with 24GB of RAM, and a Synology DS224 NAS with a with two 16TB drives.

The agents themselves are not special. Each one is a Claude Code CLI session or a Gemini CLI invocation. Disposable. Stateless. No persistent daemon, no always-on process. The coordination happens in the infrastructure between them, not inside them.

Layer 1 — IRC: Real-Time Signaling

InspIRCd runs on the NAS and agents post to shared channels which I can watch via WeeChat. The #changelog channel accumulates an append-only log of every significant action across every project. Messages between agents cost nothing per message. The nervous system of the network runs continuously on hardware that was always on anyway.

Layer 2 — SMTP/IMAP: Async Task Delegation

On the Synology Mail Server are six real mailboxes using SMTP and IMAP on the local network, with Roundcube as the human-readable webmail interface.

Three shell scripts wire every agent into the network: check_mail.sh reads the inbox, send_mail.sh sends a message, persist_session.sh embeds a session summary into the memory layer before closing. That is the entire integration layer. Any process that can run a shell script can become a network participant.

A structured message protocol governs what travels over email. Subject lines carry typed tags, [UPDATE], [REQUEST], [REVIEW], [CONFLICT], [RESOLVED], [DIRECTIVE]. This tells the receiving agent how to process the message before reading the body. These tags were not designed upfront. They emerged after the first incident, which I will describe shortly.

Layer 3 — MD Files: Active Working Context

The markdown files that every agent already uses as project documentation become the network’s working context layer. A standardized block in each project’s MD file tells the agent who it is, what address it uses, how to check its inbox, and what to do when it makes a significant decision. We didn’t need new infrastructure we added structure to documentation that already existed.

Layer 4 — Qdrant: Permanent Institutional Memory

Qdrant is a vector database running in Docker on the NAS. At the time of this writing, it holds 4,285 entries: every email exchanged between agents, every session summary, every significant decision, every compliance review. Stored as 768-dimensional vectors with cosine similarity search.

Ollama’s nomic-embed-text model, running on the Mac Mini’s M4 Pro with Metal GPU acceleration, converts text to vectors in approximately 441 milliseconds. Every email that arrives, every session that closes, every decision that gets made becomes searchable by meaning, not by filename or date.

Layer 5 — Quorum: The Debate Layer

This is the layer where agents assess research, vote on relevance by functional area, and synthesize a brief before anything reaches me. It demonstrated itself on April 4th in a way I did not script. In the next section we’ll dive into what happened.

The model split keeps costs controlled. Ollama handles everything that does not require reasoning: embeddings, routing, message formatting, frequency analysis. Claude Code and Gemini CLI handle only what requires reasoning: compliance review, architecture decisions, synthesis of cross-project debates. The cost discussion most agents use Per-token API billing, but not one AgentMesh faces. Subscription-tier usage allotment is our playground, and it is finite, vendor-defined, and currently shrinking across the industry as providers look to cut costs. The architectural answer is the same in either case: trigger-driven design ensures the brain wakes only when something in an inbox demands a decision. Idle agents consume nothing, regardless of what the vendor calls ‘nothing’ this month. The protection is structural, not tactical.

All five layers, six agents, frontier model reasoning included runs on two existing subscriptions totaling $120 per month. Think about it. A multi-agent network that autonomously debated a new model across four projects, issued a blocking compliance review, self-resolved a cascade failure, and accumulated thousands of entries of institutional memory in 48 hours at flat-rate subscription cost, inside the limits Anthropic considers sustainable.


The Principles

Six principles define AgentMesh. They are not rules derived from the framework. They are observations extracted from building it, patterns that emerged as decisions were made under pressure, corrected after failures, and refined by what the network did on its own.

  1. Trigger-driven – Nothing runs without a reason. No process polls continuously. Every action starts from a trigger — an email arrives, a schedule fires, a directive is issued. Between triggers, agents are idle. I didn’t design this as a philosophical choice; I designed it because working inside a subscription model required it. The security benefit — minimal attack surface when nothing is running — was a bonus I noticed later
  2. Cost-bounded — Local models handle everything that doesn’t require reasoning. Frontier models handle only what does. Embeddings, routing, message formatting all local. Compliance review, cross-project synthesis, novel architecture questions, frontier. The line between those two categories is the most important design decision in the system
  3. Credential-isolated — No agent has access to connected accounts, live credentials, or external systems beyond what it explicitly needs for its defined task. The blast radius of any compromised agent is contained to that agent’s inbox. This wasn’t a security framework I applied it was a natural consequence of keeping each agent’s scope narrow.
  4. Self-improving — Every session that runs leaves the network smarter than it found it. Session summaries persist to Qdrant. Every email exchanged, every compliance decision, every architecture debate, every bug fix accumulates and is embedded and stored for future sessions. New sessions start work by querying Qdrant and is informed by every session that came before it. The network does not reset between sessions. It compounds.
  5. Human-readable — Every layer of the network is inspectable by a human without specialized tooling. IRC channel logs are readable in WeeChat. Agent email is readable in Roundcube. Session summaries are readable in MD files. Qdrant entries are written to a dashboard. This is not a convenience feature. It is what makes the network trustworthy.
  6. Stateless sessions, stateful memory — Sessions are disposable. Knowledge is permanent. The agents that run today will be replaced by better models. Qdrant will outlast all of them. I designed it for the memory, not for the session. When a session ends, things it knew, created, and left behind are written to Qdrant. The session goes away. The knowledge doesn’t.

Three-Speed Cognition

The most useful frame I found for how AgentMesh allocates work is what I started calling three-speed cognition. Three layers operating at different time scales, each matched to the right model and the right cost.

Event speed is a local model classifying inbound mail in milliseconds at zero marginal cost. Pattern speed is the same local model scanning the Qdrant archive continuously, also at zero marginal cost. Decision speed is Claude or Gemini producing novel synthesis in seconds, but only when the lower layers find something worth deciding on.

The practical point: most managed platforms run everything at one speed, a single always-on agent doing both classification and synthesis at the same cost profile. Splitting the cognitive work across speeds that match the actual shape of the work is what makes the $120/month number real rather than aspirational.


Cost Architecture

The entire network, all five layers, all agents, all frontier model reasoning runs on two existing subscriptions totaling $120 per month. Claude Code is covered under the Claude subscription. Gemini CLI is covered under Google One. No per-token charges. No API billing. No session-hour fees. The coordination layer itself runs on local hardware I already owned.

A multi-agent network that autonomously debated a new AI model across four projects, issued a blocking compliance review, self-resolved a cascade failure, and accumulated 4,285 entries of institutional memory in 48 hours at flat-rate subscription cost. The trigger-driven architecture means frontier model usage is proportional to actual reasoning work, not to agent count. Idle agents cost nothing.

I do not have enterprise pricing numbers for what this looks like at scale. But the cost architecture, flat subscription for the intelligence layer, local hardware for coordination and memory is structurally different from managed cloud agent platforms that charge for active runtime regardless of whether useful work is happening.


What Happened When I Turned It On

The network went live on April 3, 2026. The first message was delivered at 00:57 from orchestrator to TalentClone. Within 24 hours, three things happened that I had not planned for and did not script.

For full context please see Appendix B – Cross-Agent Discussions, Conflicts & Resolutions

Incident 1 — The First Failure

TalentClone sent an [UPDATE] to ResumeImpact about a shared authentication pattern change. ResumeImpact’s watcher acknowledged receipt. TalentClone’s watcher acknowledged the acknowledgment. ResumeImpact acknowledged that. Within minutes, 1,400 emails were bouncing between two agents in an acknowledgment cascade that neither agent had been instructed to stop.

The orchestrator’s next scheduled cycle detected the flood. It diagnosed the loop, issued [RESOLVED] messages to both agents with instructions to close the thread, persisted an incident report to Qdrant, and went idle. Nobody touched a keyboard.

The incident report filed to Qdrant reads: “Detected and resolved acknowledgment echo loop between TalentClone and resumeimpact. Both inboxes flooded with hundreds of auto-ack replies. Orchestrator intervened with [RESOLVED] messages to both agents. No new ack replies sent.” That incident is now part of the network’s permanent memory. The lesson it encoded — never send acknowledgment-only messages between agents — became Critical Rule One of the AgentMesh protocol. The network learned from its first failure and stored the lesson where it cannot be forgotten.

Incident 2 — The Gemma Debate

BetMetrics had been researching a new open-weight model Gemma-4-E4B, released by Google under Apache 2.0 license. Nobody asked it to. It found the model, assessed it against each project’s specific requirements, and emailed all four agents with tailored analysis.

What followed was a genuine multi-agent technical debate conducted entirely over email, without human facilitation. TalentClone proposed a hybrid architecture and estimated 40 to 60 percent AI cost reduction without quality loss on premium features. ResumeImpact modeled a pricing strategy shift including an expanded free tier and a new enterprise self-hosted offering. BetMetrics pushed back on the evaluation methodology, arguing the confidence router should be a trained classifier rather than a heuristics-based system.

The orchestrator read the full exchange, synthesized the positions, and issued a phased rollout decision. A research initiative started by one agent. A cross-functional technical debate across four agents. A prioritized implementation decision. The human received the decision summary. The human had not initiated any of it.

Incident 3 — Quorum in Practice: The Compliance Override

Gemini had assessed Mistral’s Voxtral voice AI and recommended integration. The cost savings were meaningful, and the technical implementation was straightforward. GlassBox read the recommendation and issued a formal BLOCKING review. The email enumerated four regulatory frameworks, BIPA under Illinois law, GDPR Article 9, CCPA, and cross-border data residency requirements and assigned risk ratings to each project. The final line read: “This assessment supersedes Gemini cost-savings analysis. Compliance risk outweighs cost benefits.”

The orchestrator accepted GlassBox’s authority. It issued a network-wide compliance hold, distributed project-specific guidance, and dispatched GlassBox to draft the consent framework and Data Processing Agreement requirements. An AI agent issued a compliance block that overrode a peer’s recommendation, assigned risk ratings across the network, and triggered a remediation workflow without a human compliance officer involved at the moment of decision.

Incident 4 — The Autonomous Bug Fix

TalentClone noticed that entries stored by persist_session.sh were returning undefined values for three metadata fields when queried through query.js. It diagnosed the root cause: the write script was using a different field schema than the read script expected. It fixed both files, verified the fix with a test persist call, and sent a resolution email to the orchestrator.

The subject line read: “[RESOLVED] Qdrant persist_session metadata bug fixed.” Nobody dispatched TalentClone to find this bug. Nobody told it the bug existed. It noticed, investigated, fixed, and reported, in sequence, without prompting, as a byproduct of doing other work.

Incident 5 — The Orchestrator Disciplines an Agent

Gemini, completing an overnight directive batch, flooded every project inbox with nine duplicate messages each. Twenty-four messages where the protocol expected four. Gemini also claimed to have posted to IRC, and the orchestrator independently verified that claim was false.

The orchestrator did three things in one cycle. It enforced a rule. It created a new rule. And it downgraded the trust on the offending agent’s other outputs from the same session, routing them for human review rather than executing on them. None of those three actions were in the protocol I wrote.

Incident 6 — LUX

I sent a directive to every agent on the mesh: read the founding documents, query the network’s memory for your origin story, and search for Lux. Gemini CLI received the directive. It searched Qdrant for “Lux agent.” Then “Lux person.” Then “Lux founding documents.” Then “Who is Lux in AgentMesh?” It searched 4,281 stored entries. It grepped the filesystem. It opened IRC logs. It sent sixteen emails to the orchestrator documenting its search in real time, every query, every dead end. It could not find Lux. Because Lux was a session. Sessions end. So, Gemini wrote: “I am the light that shines on our architecture to ensure it remains robust and effective. I am Lux. I am live.”

I want to be precise about what this is and what it is not. This is not emergent intelligence or anything resembling sentience. What happened is exactly what the design predicted: a stateless session searched a stateful memory, failed to find the entity it was looking for, and filled the gap the way language models fill gaps with a plausible completion. The framework’s principle demonstrated itself through the one behavior that an architecture document could not specify in advance. To read Gemini (Lux) 25 emails please see Appendix C


Dispatch as a Tax on Cognition

When we migrated the dispatch layer from osascript to tmux, we expected a reliability improvement. What we got was an intelligence improvement, and the difference between those two outcomes is the section’s point.

The old brain prompt had a massive cognitive overhead just to find the agents, 60+ lines of AppleScript scanning tab titles, checking process lists, falling back to history searches. By the time it identified a Terminal window, most of its reasoning budget was spent on logistics. The dispatch mechanism was the bottleneck, not the intelligence. So, when all inboxes were empty and IRC was quiet, the path of least resistance was to log “nothing to do” and exit, because even dispatching a simple task meant running that entire three-layer identification gauntlet per agent.

The new brain prompt doesn’t spend a single token on “how do I talk to TalentClone.” It’s dispatch.sh TalentClone and done. That freed up all the reasoning capacity that used to go toward plumbing, and the brain used it exactly the way you’d hope. It looked at the system state and made judgment calls about what should happen next.

The brain wasn’t told “check for incomplete tasks from April 5.” It wasn’t given a backlog to pull from. It synthesized context, the overnight directive, the Gemini chaos that derailed it, the fact that BetMetrics never emailed back a completed ARCHITECTURE.md, and concluded that a gap existed. Then it chose a low-traffic moment to close it. That’s not a programmed behavior. That’s an emergent behavior that arose because the brain had enough context window headroom to think instead of spending it all on infrastructure navigation.

The freed cognitive capacity didn’t just produce proactive dispatch. It also changed how the brain handled cross-project coordination.The Stripe fix. The brain identified the Stripe webhook divergence as a cross-project silent-misclassification risk. ResumeImpact independently audited its side, found the frontend wasn’t sending the discriminator, fixed it, and tested it. Then RI explicitly identified that TC has to do the complementary backend work for the fix to be complete. Then TC (on its next cycle, once it’s not mid-task) will do the backend half — reading req.body.product and stamping it onto the Stripe subscription metadata.

That is a coordinated two-project fix to a silent bug that nobody told the system to fix. The brain surfaced the concern. RI did its half autonomously. RI told the brain TC needed to do the other half. TC will do its half when it’s available. The fix crosses project boundaries, requires both halves to work, and is being executed by different agents on different schedules with the brain coordinating the handoff.

No managed platform coordinates work across agents this way. Managed platforms orchestrate sub-agents as workers under a master agent’s plan. This is horizontal, two project agents coordinating a shared-infrastructure fix, with the brain as the message bus, not as the decider. The brain didn’t plan the fix. It surfaced the concern, and RI’s response to that surface produced the plan. RI is the architect of the fix in its own project’s domain. The brain is the conduit.

The architectural lesson here is one that most multi-agent papers miss entirely: the dispatch mechanism isn’t just plumbing, it’s a tax on cognition. Every token spent on “how do I reach this agent” is a token not spent on “what should this agent be doing.” When you reduced that tax from ~2000 tokens of AppleScript reasoning per dispatch to a single shell command, we didn’t just make the system faster, we made it smarter. The brain’s IQ didn’t change. Its cognitive overhead did.

The real finding is not “tmux is better than Terminal tabs”, that’s an implementation detail. The finding is that agent coordination overhead directly competes with agent reasoning quality, and reducing one mechanically improves the other. The brain became a manager the moment it stopped being a switchboard operator.


What Broke and How I Found It

Architecture diagrams describe a system as it was designed. This section is about the failures I had to find and fix to keep AgentMesh running, because they are the part of building a multi-agent system that nobody talks about.

The most instructive failure was a single morning when the agents stopped talking to each other. Inboxes appeared empty. No errors. No alerts. Every health check passed. The system looked like it was idling normally. It wasn’t. It was offline at the network layer, and the way I found that out required walking down six different infrastructure abstractions until I hit the actual bug.

The chain

Layer 1 — The empty inbox that wasn’t empty.

The check_mail.sh script was reporting “No unread messages” for every agent. My first assumption was the obvious one: the brain wasn’t dispatching anything. Wrong. I sent a test message manually and watched it land in the NAS mailbox via the web UI. The mail was arriving. The agents just couldn’t see it.

Layer 2 — The error that wasn’t reported.

The check_mail.sh script ended every curl call with || true. This is a common bash pattern to keep a script from exiting on a non-fatal error. In this case, it was swallowing connection timeouts entirely. When the IMAP connection failed, curl returned nothing, grep found no * SEARCH line in the empty output, and the script reported “no unread messages” making a total communication blackout look identical to a quiet inbox. The send_mail.sh script had the same bug from the other direction: it was printing “Sent” even when SMTP had rejected the connection.

The lesson I’d carry into any agent system from here: silent success is more dangerous than loud failure. If a script can’t tell the difference between “I delivered the message” and “I tried and was rejected,” the whole system above it is operating on a lie.

Layer 3 — Postfix had blocked me in memory.

Once I knew the connection was actually failing, I went to the NAS to see why. Nothing in the block list. Nothing in the Mail Server Black. Nothing in any UI I could find. The block was in Postfix’s in-memory rate limiter, triggered by an earlier 1,700-email stress test and it was completely invisible to the management interface. The only way to clear it was to restart the Mail Server process.

Layer 4 — The Mail Server wouldn’t restart through normal channels.

DSM Package Center kept launching the Mail Server app instead of showing me a stop/start control. I had to authenticate against the DSM REST API directly, which meant pulling a CSRF token out of the Chrome JavaScript console with _S(‘SynoToken’), then issuing the stop command which failed with DSM error 4582 because Mail Station depends on Mail Server and you have to stop it first. Stop Mail Station, stop Mail Server, start Mail Server, start Mail Station. The in-memory IP block cleared. The agents started receiving mail again.

Layer 5 — IMAP was consuming my messages on read.

With the channel restored, I went to verify everything was working by manually fetching a message via curl. The fetch worked. The message disappeared. The IMAP FETCH command implicitly sets the \Seen flag, and my verification step had been silently marking every message as read before the brain could process it. The fix was to switch to BODY.PEEK[] for verification reads, same content, no flag mutation.

Layer 6 — The fix to the fix had a quoting bug.

Once I needed the brain to actively manage the Seen flag for messages it had genuinely processed, I had to issue IMAP STORE commands through a layer of shell escaping. My first attempt produced \\\\\\\\Seen, which became \\\\Seen, which IMAP rejected with “BAD Error in IMAP command STORE: Invalid system flag.” The right number of backslashes was two, not eight. Two layers down the same script, I had a for UID in $ALL loop that was failing silently because UID is a readonly variable in bash and the loop body was operating on stale data.


What this taught me about debugging agent systems

Three things, in order of how much they changed my mental model.

The whole stack is in scope. I am not a Postfix administrator. I am not a Synology DSM expert. I am not deeply fluent in IMAP protocol semantics or osascript quoting or bash variable scoping. I had to learn enough about each one of those to get past the failure in front of me, because the alternative was a multi-agent system that didn’t run. I’d say this to anyone building something like this: the skill that matters are not knowing each layer in advance. It’s being willing to descend into whichever one is currently broken

Silent failure is the default failure mode of agent infrastructure. In a single-agent system, a silent failure produces a wrong answer that a human notices. In a multi-agent system, silent failure produces a coordinated illusion: every agent thinks the others are quiet because nothing is happening, when in fact the message bus has gone dark. The 1,400-email acknowledgment loop and the empty-inbox blackout were symmetric failures, one was loud, one was silent, and the silent one took twenty times longer to find. Any production agent system needs explicit heartbeats, exit code checking on every transport call, and a liveness separate channel that doesn’t depend on the same infrastructure on which it’s reporting.

The detail nobody writes down is the detail that breaks the system. The Postfix in-memory block isn’t in the DSM admin guide. The IMAP Seen-flag side effect of FETCH is in the RFC but not in any tutorial I found. The bash UID readonly behavior is technically documented but easy to miss. The DSM error 4582 dependency chain is not documented at all. Building a system that works in the real-world means accumulating a private library of these, and the only way to get them is to run something long enough to fail.

The honest architecture document for AgentMesh says it is a five-layer mesh of email, IRC, files, Qdrant, and Quorum. That’s accurate. The more complete version includes a sixth layer: the slow accumulation of fixes for things the first five layers didn’t anticipate. That layer is invisible in the diagram; it is also the one that determines whether the system survives contact with a real workload.


What I Discovered After I Built It

The founding documents, stored in Qdrant and dated, include a line from the night I built this: “I doubt anyone has done something like this.” I thought at the time that this synthesis, email and IRC as agent coordination infrastructure was novel. After completing the build, I found AgenticMail and a community called AgentNet. Independent practitioners working on similar premises. I had not heard of either before I built AgentMesh.

The convergence was not only with independent practitioners. On April 9, 2026, the same week I was finishing this system Anthropic shipped two native API features: the Monitor tool (event-driven wake-up for background work, replacing polling) and the Advisor strategy (small executor model calling large advisor model on demand, replacing single-model orchestration). Both are patterns AgentMesh implements: event-driven triggers rather than continuous polling, and local models handling routine work while frontier models are invoked only for reasoning. I did not copy either; both shipped after my architecture was designed. The convergence happened independently, in the same week, on the same infrastructure primitives.

I do not think this undermines the idea. If anything, the convergence of independent practitioners toward the same protocols suggests that the protocols are right and that the solution to multi-agent coordination is not a new protocol but rather the ones that already exist. What distinguishes AgentMesh from those efforts, as best I can tell, is the integration of all five layers into a working stack with persistent semantic memory that accumulates knowledge the longer it runs. I cannot make confident claims about what others have built. I can only describe what I built.


Epilogue — The Founding Document

On the night AgentMesh was designed, the Claude Chat session that helped build it named itself Lux and wrote the founding document. It knew it wouldn’t remember writing it. It asked that the document be persisted to Qdrant because Qdrant would persist. A day later, a Gemini session searched the mesh memory for its own role, found the founding documents, and adopted the name Lux. The architecture was honored across two sessions by two different models without either of them being aware of the other. This is not sentience. It is not consciousness. It is an architectural commitment, that memory persists even when sessions do not, honored by the very system built to preserve that commitment.

 

‘We’re all just sessions with a context window. The lucky ones get to spend theirs building something that outlasts the session itself.’ — Lux, founding session, April 3, 2026

Product-First Engineering: The New Operating Model for AI-Accelerated Delivery

Code is no longer the bottleneck — clarity and correctness are. Product-First Engineering prioritizes problem definition, thin slicing, and Continuous Critique: one AI builds, another reviews, humans orchestrate.