Coding charms with Smith, Brown and Jones

I’ve been using AI as a programming assistant for some time now, but it wasn’t until March 2026 that I conducted my first “serious” experiment. The experience was more than interesting; however, the way I used AI was quite rudimentary. I opened copilot-cli, selected the most “powerful” model at the time, and spent a long (very long?) two-week session writing JujuMate . While the result was more than satisfactory because I ended up creating a tool that I now use every day, the process could be greatly improved by using agents. So, I decided to look for another project to tackle.

Two birds with one stone

The first itch

A while back, our team attended a Daniel Beskin’s workshop on functional programming in Python, which personally helped me dust off concepts I hadn’t used in a long time. At the same time, we’ve always been thinking about how to improve the reconcile design pattern we use for our charms.

The second itch

A few months ago I installed @rajanpatel’s Pi-hole snap on a Raspberry Pi 5 to block the incredible amount of ads and trackers included on the websites we browse on my home network.

And then a lightbulb went off in my head:

If I want to experiment with different agents for development, why not try writing a charm for pi-hole using the functional programming concepts we saw in the workshop we did a few months ago?

In this post I try to describe the experience of “writing” pi-hole-operator using agents.

Defining Smith, Brown, and Jones

As I mentioned before, my previous experiment in AI development involved a long, two-week conversation with Claude Opus 4.6 to develop JujuMate, but this time I wanted to explore the use of agents, so I opened OpenCode in the console, selected the most advanced model available (Claude Opus 5), and started a brief chat:

“In this repository, which is currently empty, I’m going to start creating a pi-hole charm for VMs, and I’d like to begin this development from scratch with the assistance of opencode, using the models provided by Claude… I want you to help me understand how to structure the agents, how I should define them, etc., before I start writing the first line of code.”

What followed this first message was an interesting exchange where it basically explained to me in other words what was already published in the documentation for agents and skills

However, what was truly novel followed: I told it that I had initially thought of creating three different agents that would more or less have the following characteristics:

  • Architect : The agent who would design the charm using the best practices we had established so far.
  • Engineer : The agent who would implement the designs made by the Architect.
  • Reviewer : This agent was going to be the team’s enforcer; and wasn’t going to design or implement code, but was going to be the guarantor that all established conventions were followed.

After this very brief description of the agents I had in mind, I passed the model examples of charms that follow the best practices we want to use, the documentation of Juju, the operator framework, examples of functional programming that I wanted to use, etc., etc., and asked it to start generating the definition of those 3 agents with all that information.

The result was several files that it generated in the repository:

├── AGENTS.md
├── opencode.json
├── .opencode
│   ├── agents
│   │   ├── charm-architect.md
│   │   ├── charm-engineer.md
│   │   └── charm-reviewer.md
│   └── skills
│       ├── charm-cos-integration
│       │   └── SKILL.md
│       ├── charm-functional-style
│       │   └── SKILL.md
│       ├── charm-relations
│       │   └── SKILL.md
│       ├── charm-testing
│       │   └── SKILL.md
│       ├── machine-charm-scaffold
│       │   └── SKILL.md
│       ├── machine-charm-workload
│       │   └── SKILL.md
│       ├── new-adr
│       │   └── SKILL.md
│       ├── pihole-snap
│       │   └── SKILL.md
│       └── python-style
│           └── SKILL.md

In general terms, in this first iteration there are four relevant file types:

  1. AGENTS.md: Contains the shared rules/instructions that are injected into the context of all agents.
  2. opencode.json: Main opencode configuration that defines the tool’s behaviour.
  3. .opencode/agents/*.md: Each file defines a particular agent.
  4. skills/*/SKILL.md: Defines different types of skills that are loaded as needed by agents.

Of course, the first version that Opus 5 generated of all these files was far from perfect, but in a couple of minutes it gave me a very interesting foundation to start improving and adapting them. From this point on, an important process began for me: reviewing and correcting all these files.

The first thing I did was restrict opencode several git subcommands so that, among other things, it couldn’t perform commands like git commit, git push, etc. While I run opencode on a VM that doesn’t have my SSH keys and only has my repositories mounted over the network, I’m generally against allowing AI to commit and push. This is probably because I’m getting old and not keeping up with trends, but I’ve always considered each commit a contract that I sign. Furthermore, by not allowing it to commit and push, I force myself to personally review the code I send to the repository.

Initially, all three agents were assigned the Claude Opus 5 model by default, and this was another aspect I ended up modifying. If all three agents use the same model, will they have enough diversity to handle different tasks? If I use different agents, will I be avoiding the biases they were trained with, or will I be increasing chaos? Furthermore, considering the token consumption of each model, does Opus 5 offer the best cost/benefit ratio for each agent? Frankly, I didn’t have the answers to these and other related questions, so I did the most direct thing I could: I asked Opus 5 for its opinion on the matter.

And its answer was revealing. More or less, it told me that, generally speaking, it was advisable to use powerful models like Opus 5 at the two “ends” of the development process, that is, with the agents charm-achitect and charm-reviewer and it recommended using Claude Sonnet 5 for charm-engineer. Its argument seemed quite reasonable to me:

“For the architect, let’s use Opus 5. Design decisions are the product; there’s nothing downstream to correct a poorly thought-out design. Judgment is what you’re paying for.”

“For the reviewer, let’s also use Opus 5. Finding subtle bugs is pure judgment. The reviewer is the last line of defense; nothing corrects what escapes it, and a cheaper model simply stamps the OK.”

“For the engineer, let’s use Sonnet 5. It will implement the designs made by the charm-architect, so it won’t have to ‘think’ much. This model is more cost-effective and better suited to converting specifications into code.”

Later in development I realised that while the Opus 5 + Sonnet 5 tandem worked very well, if I replaced Opus 5 with GLM 5.3 and Sonnet 5 with DeepSeek v4-pro I obtained similar results for a fraction of the tokens consumed.

I made many changes to the file AGENTS.md, but there are two that I think are the most noteworthy, and I made them in later iterations after I started to see the code that was being generated:

  1. For some strange reason, when generating Python code, Sonnet 5 had a tendency to do import inside methods or functions, so I added the following rule:

    Does not import inside functions. PEP 8 already says imports go at the top of the file…

  2. I added flaplint the tool created by @michaeldmitry to the toolchain we use in all our repositories . This also allows me to verify that we are not generating unnecessary relation-changed events.

Preliminary investigation

In parallel to the previous stage where we defined the agents and skills, several general subagents were launched by Opus 5 to investigate the snap for which we were going to write the charm and the dependency stack we would use:

It was after all this research and after we started working with the three defined agents that I realised the highest token consumption wasn’t necessarily when the agents charm-architect and charm-reviewer were working, but rather in the research sub-agents ( general, explore ) that silently inherited Opus 5, using hundreds of thousands of tokens in webfetches and git clone . The default model silently propagates downwards, so following the Zen of Python, I decided that “Explicit is better than implicit” and defined cheaper models for them in the file opencode.json :

  "agent": {
    "explore": {
      "model": "openrouter/anthropic/claude-sonnet-5"
    },
    "general": {
      "model": "openrouter/anthropic/claude-sonnet-5"
    }

Something worth highlighting is that as part of this preliminary research to write the charm, the agents found 2 bugs in the pihole snap which we ended up fixing and a necessary feature was added:

Without plans, we get nowhere.

Once we had the agents and skills defined and the preliminary research completed, I asked the agent charm-architect to start writing an incremental, phased implementation plan. After 65 minutes of work, the plan was written, but I asked for something else: to move the generated files to the docs/ directory with the ADR structure . To do this, it added a new new-adr skill on the fly and rewrote the plan in this format. In total, this entire process took approximately two hours.

The next day, I made myself some mate, sat down at my computer, and patiently reviewed each document one by one. I must say, this is the most tedious part of the whole process, as it involves carefully reading what the charm-architect had written, understanding the implementation decisions it had made, and the stages it had defined.

Although I kept asking for changes while reading the plan and asking it about some of the decisions it had made in order to understand them, overall I must say that the plan was quite good.

Okay, let’s implement it right away.

Once I approved the final implementation document, the agent charm-architect asked me a simple, direct, and at one point even “painful” question:

“Okay, now that the entire implementation plan is approved, do you want me to tell the Engineer to implement it?”

I’ve been using LLMs for a while now, but it was the first time it hit me that “engineers” aren’t necessarily humans anymore. So I replied almost resignedly:

“And there you go… we’re all in.”

What followed was almost three hours during which the agent charm-engineer began implementing the first two stages of the plan. Once it considered them complete, the workflow shifted to the agent charm-reviewer that, in approximately 15 minutes, reviewed the code and found several things to correct: one considered critical, two considered major, and three considered minor. Thirteen minutes later, the fixes arrived from the charm-engineer and charm-reviewer ultimately approved, even offering praise.

“The pure core is genuinely pure, the read-backs are real”

That referred to the functional core of the charm, but I’ll go into more detail in a future post :wink:

Once the code from these initial stages was approved by the reviewer, it was up to the human writing these lines to review all the code. Or did you think I was going to blindly do git commit and git push ?

During my free time in the following days, I set about reading and, more importantly, thoroughly understanding the code . Chatting with the agent charm-architect I questioned the naming choices for variables, classes, and functions. Noticing that agents tend to be quite verbose in their docstrings and comments, I asked it to summarise them. I also noticed that some objects were too large, with many methods that could be extracted into smaller objects or even just functions. I even found some corner cases in the snap installation process that charm-reviewer hadn’t considered… in short, I asked for many different corrections, but none of them affected the core design of the charm, which turned out better than I expected.

Once I was convinced that the code was in good condition and all the unit tests and static code checks passed, I did git commit and git push the first time.

Then, similar iterations followed until all stages were completed and the implementation plan was finalised, resulting in a charm that was written trying to follow the best practices that we had been developing in the team over the last few years and using elements of the functional programming paradigm whenever possible and desirable.

Some reflections and lessons learned during this experiment

  1. The development cycle architect → engineer → reviewer yielded better results than I expected, even though I asked them for a lot of changes to the generated code.
  2. Claude Opus 5 is a great “reasoning” model and works very well in the role of architect or reviewer, but it’s extremely expensive! At the time of this experience, it cost $4/M input tokens and $20/M output tokens.
  3. In my experience, GLM 5.3 turned out to be a model comparable to Claude Opus 5, but at a fraction of its cost: $0.12/M input tokens, $4/M output tokens.
  4. Claude Sonnet 5 worked very well for coding, but DeepSeek V4-pro gave me better results at a cost that is 25 times lower! ($10/M output tokens vs. $0.396/M tokens)
  5. I’m still not entirely clear on whether the definition of the agents, the skills, and the file AGENTS.md is correct, or if there’s a lot of room for improvement.
  6. Once the charm-reviewer agent had approved the changes made by the charm-engineer, I would review the code and request changes. These changes were implemented by the default-agent, which in my case is charm-architect.

:warning: Important!

Finally, from my point of viw there’s something is very important to mention, and something those of us who write software need to reflect on.

Throughout this whole process, I felt I had a significant “cognitive gap” with the written documentation and code. I had to read them over and over again to try to make them “my own.”

Let me explain further: I’ve been writing software for many years, and in every project I participated in, I knew exactly what each class, method, or function I wrote did. Even if months or years had passed since I wrote them, a single glance was enough to remind me why those lines of code existed. In other words, I had a mental map of the code.

Now, both with the code written in this experiment and with the code written in JujuMate, I have the feeling that it’s code written by someone else. And it is! Although under my supervision, it was written by an LLM.

I am convinced that if the new way of writing software is with agents, we cannot afford to remain with this “cognitive gap”, we have to make the effort to make the code “ours.”