A few minutes into the second AI self-assessment lunch-and-learn, I told the group this session was one of the committed actions in my year-long goal.

I set the goal because AI maturity is not something I can cram. It is a muscle I have to train deliberately, with evidence, in public. The lunch-and-learn is where I do that training out loud. I bring a few Improvers together, we run our assessments, we share where we are struggling, and we show the work behind the score. Some of us journal, some build skills, some run spec-driven experiments. All of us agree on one thing: it is much easier to keep a practice alive when someone else is watching.

This is a focus group, not a polished presentation. People drop in, share a screen, argue with their own assessment numbers, and admit when a prompt fell apart under pressure. That is the point. I wanted a venue where the struggle is the curriculum. If I only talk about AI wins, I learn very little. If I talk about the misses, the bad prompts, the agent that stubbed a function and faked a test, I learn what to fix next.


Why write about the lunch-and-learn? Two reasons.

First, accountability. When I say my goal out loud in front of others, I am much more likely to follow through. The lunch-and-learn forces me to show up with something: a new skill, a new workflow, a dashboard, or a hard number from the month. Writing about it afterwards keeps the promise visible.

Second, it normalizes the work for everyone else. AI maturity is a set of habits that look messy from the inside. I want other people to hear the questions we ask, the experiments we share, and the workarounds we admit. Those moments make the journey feel possible. They also make it feel less lonely.


My last assessment flagged two things: evaluation and team fluency. Since then, I did work on both.

On evaluation, I started by leaning on an LLM to evaluate my write-e2e-tests skill. The results were hit or miss, but they helped me decompose a skill into several smaller ones. One writes the outside of an end-to-end test, one implements the mechanics, and another troubleshoots failures. At the time, I was using Cypress, and the pieces started to fit better once I broke them apart.

Then I followed the evaluation-driven development module in Improving’s self-directed learning repo. Our Houston AI roundtable ran weekly katas on it. I got stuck on the PRD exercise until a teammate nudged me to try a different angle. That nudge moved me from trying to force one exercise to building a skill that writes user stories from a problem statement and turns those stories into the prompts I actually use. I use that skill every day. It works with the free SWE models, including SWE 1.7, and it gets me from a description to good stories. If I need a stronger model for implementation planning, I can escalate from there.


I also built what I am calling my digital brain (still pending a better name for it) for running these assessments. It pulls from my Obsidian vault: journals, blog posts, sprint docs, talk transcripts, skills, and workflows. Then it pushes the content into pgvector, the Postgres vector extension, using LM Studio to create the embeddings. The assessment itself runs with Gemini. I built the visualizer by dumping the raw output into Gemini Canvas, downloading the HTML, and having Devin implement it as a custom app on my desktop. The result shows a radar graph of the dimensions, the evidence behind each score, a what-if simulator, and a recommended action plan. It also lets me dig into the questions and see what the model used as proof. I do not trust the 77 or the “readiness for Stage 5+” it gave me yet (it is still too nice), but I can now see exactly where it is being generous and update the ingestion or the prompt.

Pasted image 20260804104040.png

The same system runs other assessments: my skills that are listed in our internal platform, my job experience, even my StrengthsFinder profile. Each one links back to real artifacts. I also feed the assessment into NotebookLM so I can ask follow-up questions and get mind maps or podcast-style explanations of the recommended next steps. That is the kind of evidence I want in front of the group. Not a claim. A chart. Not a prediction. A number I can explain.


On team fluency, I made the repeatable multi-agent workflow an explicit part of the assessment. The team has to be able to leverage what I build, or the work is just mine. I built the workflow so the alignment is part of the process, not an afterthought.


The cost story is just as concrete. In June, I was 111% over my AI budget.

Pasted image 20260804222938.png

In July I switched to the free SWE models and stopped using paid M365 Copilot. My AI spending dropped about 65% compared with the previous period. It would have dropped 77% if I had cancelled my Copilot seat earlier.

Pasted image 20260804223140.png

The dashboard seems to indicate I almost didn’t use AI since July 10, but that’s because it shows ACUs…

Pasted image 20260804223046.png

I track cost against message volume on a dashboard, and the relationship is clear: as I got better at choosing the right model for the task, my messages went up and my cost went down. The tool did not get cheaper. I got better at using it.

Pasted image 20260804223333.png The purple circle on the far bottom right is SWE 1.7, and the one to its left is SWE 1.6.


I am doing more with AI now than I have in the last two years: client work, Improving’s internal work, the Engage platform, thought leadership, and journaling that turns into blog posts. The token cost is essentially free. That includes writing and editing these blog posts daily with SWE and the skills I have built.


I also write it down because memory is fragile. We talk about many things in an hour. One of us might mention a new skill. Another might describe a workaround. I will try a different evaluator. If we do not capture those threads, they disappear. The post becomes a shareable artifact we can search, reference, and improve. It is part of the system, not just a meeting.

The session ended with a proposal to collect our skills, evaluators, and prompt improvements in one place. That idea came from the same place as the lunch-and-learn itself: if we all keep our experiments private, we each pay the full learning tax. If we share them, we split it.


Showing my work before I feel ready is worth more than the published version.


This post is part of a longer thread on how I am measuring and improving my AI maturity. Earlier posts in the thread:

Pasted image 20260810191125.png

Leave a Reply

Trending

Discover more from Claudio Lassala's Blog

Subscribe now to keep reading and get access to the full archive.

Continue reading