fifthrevision

  • Writing
  • unfiltered
  • Bookshelf
  • Projects
  • About

unfiltered

From my head to yours, no filter.

  • October 2026
  • September 2026
  • August 2026
  • July 2026
  • June 2026
  • January 2026
  • November 2025
  • June 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025

Goodbye Jekyll, hello Astro

The first commit of this version of the website happened in February 2017 (169e91e7). It went from a set of handcrafted HTML and CSS files (edit: I was wrong, it was a Wordpress site!) to a site statically generated with Jekyll.

Jekyll gave me the vocabulary and the confidence to do bigger things. The loose collection of pages turned into a coherent set of essays and portfolio entries. Then, a reading list and reviews I would write of them (that ended up being a tough thing to keep up on). Later, a microblog.

In the last few years, I’ve noticed creakiness in the Jekyll ecosystem. Once, I upgraded the Ruby version and the build broke. Plugins stopped getting maintained (I’m looking at you, jekyll-paginate-v2). Liquid 4.x is limited in what it can express (5.x looks like it’ll never get supported). Lastly, I’m not a Ruby person and I don’t want to maintain custom extensions.

Recently, I’ve used Astro for a few content-heavy static sites and I liked what I saw. Well, first and foremost, I’m a TypeScript person. My requirements for this site remain the same: static-site generation primary, JavaScript as an enhancement layer, and building on top of a variety of content types. Back then, I flirted with the idea of porting the site to Astro, but the idea of undoing years of technical debt felt daunting so I never took the next step.

Yesterday, I bit the bullet and ran the migration using Claude Code. It was done in a few hours with Fable 5.1 driving Opus 5.5 subagents and the entire process went as smoothly as one can hope. Layouts went from Liquid to Astro components. The content folders became more sane; it’s interesting how many Jekyll conventions can quickly become limitations once the use cases outgrow what they were designed for. I took the opportunity to convert SCSS to plain old CSS — CSS is mostly at parity these days with a few features in the pipe. The codebase reads better, runs better (no more fumbling with bundler, why no run scripts?), and the generated output is basically byte-identical.

Then, I shipped my first site change in years, a revamped projects list. Which is to say, fifthrevision is now officially on Astro. So far so good! I’m sure I’ve traded one set of constraints for another. Time will tell.

For nearly a decade, you served me well. Thank you, Jekyll.

Posted Oct 06, 2026

Anti-social social technology

The first question we ask of a new technology is what can this product do for me. It’s a useful lens for evaluation, and AI is right in the middle of answering that: do my work, read my emails, write my docs, achieve my dreams, yada yada.

The question seldom asked is what can this product do for the other. How do we make it easier for the other person to understand us? How do we take the burden off other people that we work with? How do we help other people? Much of what we use today don’t arrive at this stage of questioning.

AI is poised to be the most transformative technology in my lifetime. It is also, so far, a Randian ideal: it accentuates differences and reinforces individual preferences. However, we are social creatures who rise and fall together. So the imperative question we should all be asking at the moment, designer or not, is what can this technology do for us.

Posted Sep 30, 2026

This week: Axle 0.31, 0.32

I haven’t posted an update in a while because I took a vacation and had to work for money! That said, Axle shipped a few minor versions:

  • Axle 0.31.0 shipped a major CLI revamp. This update took everything I’ve learned building apps with Axle and put it into a modern, usable TUI. In addition to the new UI, Axle CLI now supports follow up interactions and resumable sessions, as well as more sensible configuration and setup.
  • Axle 0.32.0 reworked the thinking Turn and stream objects. It’s a minor but breaking change to the format. The implementation of thinking has always been a little janky, and the recent release of Muse Spark 1.3 gave me a chance to revisit it so that it’s a little cleaner.

I’m personally pretty pleased with the CLI updates. This resolved a few longstanding issues I had with the previous implementation, including the fact that it was just really ugly. With the open weight models being quite cheap and good these days – I’m really liking GLM 5.3 Flash – I’m harboring a few automation fantasies on my life. Not quite there yet, a few more gaps to close, but I’m hoping to have some updates soon.

Posted Sep 15, 2026

Intelligence and behavior

There has been constant fighting with the latest crop of AI models.

A few things I’ve noticed: AI models are much more eager to jump in and do instead of understand what you want. When they write, the output is jargon-filled and convoluted. Worst of all, the recent Opus 5 just flat out make things up; over the last few weeks, I’ve had multiple incidents where after I push on a specific point in a plan, the model tells me that it made a mistake and acquiesces.

While none of this is indicative of intelligence per se, these behaviors make the model unpleasant to use. It feels as if they don’t understand my goals, try to paper over subpar decisions, and the most egregious of all, lie. I know I’m using people words here, ascribing intentionality to floating point numbers computed across arrays of special-purpose silicon. However, language is inherently anthropomorphizing; it is what we use to express and communicate with each other, and now with computers.

The thing is, none of this is inherent to the mathematical properties that create artificial intelligence. This is all trained behavior. My guess is that the pursuit of AI models that perform well on benchmarks and long-horizon tasks that leads to choosing verifiable rewards over human feedback creates behaviors that look like this. Inattentive and dismissive.

I’ve long thought that intelligence is not a linear measure. I had a conversation recently with a friend about this where I argued that preferences–for example in language and art–are not something that can be captured by scaling laws. He disagrees. I was arguing that there are different dimensions to intelligence and maxxximizing one will diminish others. A better way to frame it would be to separate what it can do versus how it behaves.

Perhaps soon, people will start talking about IQ vs EQ again. Perhaps one day AI model makers will find the right balance between the know-it-all and the we-get-you. But I’m not holding my breath.

Posted Aug 26, 2026

This week: dtg and company

I spent a good part of this week working on a color system to token tool called dtg. It’s a somewhat niche tool that solves a long standing problem that I have. I’ll have more to say about it shortly, but suffice to say, I’m real stoked about it.

Axle got a 0.30.2 bump. Bug fix and a refinement to the prompt compactor API. It feels pretty good now. I’ve been test driving axle-code on the new stealth model on open router, ox alpha, and I’m overall really happy with how rock solid Axle is. Occasionally, I wished that I can compile typescript down to binary (yes I know about scriptc) just to really get access to real multithreading, but as it stands everything is fast enough.

Speaking of ox alpha, rumor has that it’s a GLM or GLM-based fine tune. It definitely has some Claude-isms in the text it emits (and less so GPT-isms). After running it for a bit, I’ll pin it at around a below Sonnet, above Haiku level of intelligence. In my subjective measure, coding ability is be inversely proportional to the amount of frustration I feel while using it. This new model often runs itself into a quagmire and needs me or a better model to bail it out; this is the first time I’ve had to revert my git history while using an AI agent.

I’m making little progress on the eval work that I’ve been mulling on. The approach that I’m taking — a test harness vocabulary built on knobs on LLM as a judge – feels like a dead end. It all still feels like syntactic sugar and not a meaningful step forward. Part of it is that it’s really hard to know what correct means over a long horizon, the other is that even if I do, I don’t know if I could trust it without seeing it.

Posted Aug 23, 2026

This week: Axle 0.30, new docs; Sunnyday gains skills

Axle 0.30 is a big milestone in a few ways:

  1. It marks the 30th release since I’ve started publishing the library consistently since June of last year. I typically release when I need to integrate changes into my core projects, and minor releases are when I make significant additions or breaking changes. That’s quite a pace, if I may say so myself.
  2. More importantly, the compaction API has arrived at a good place. 0.30 solved a design conundrum I’ve been facing for a while, that is, what is the role of History — the dual objects of Messages and Turns — when it can be modified by an external actor. The answer is by removing Turn, the presentation layer, out of the Agent, and reduce the surface area of data that Agents work with. Now, the data model is coherent with a bit of conceptual tradeoff which we can bridge with built-in functions.

To celebrate, and to make sure I’m explaining the conceptual model well, the documentation site got a content revamp.


Sunnyday shipped skills. It’s something I’ve been wanting to do for a while, and having a concrete use case finally gave me a reason to built it.

The implementation went through a couple of cycles. The starting point was simple enough: it walks like a file and quacks like a file. But it soon became apparent that skills do indeed have a special status in an agent platform. I’ve settled on treating skills as files in the UI while handling them differently on the backend, which let’s me keep the UI and entry points simple. Though I doubt this is the last time I’ll have to think about it.

Posted Aug 16, 2026

This week: Sunnyday Sandbox resume, Assistant; next Axle CLI?

On the Sunnyday side:

  • Agent sandboxes per session are fully resumable now, up to a maximum of 30 days. This sets up each agent up to take follow up commands, which will open up new avenues of interacting and correcting remote agents
  • Assistant is now in beta. You can now prototype, create, and update Agents using the Assistant. This is something I’ve been mulling on for a while; as much as I think the UI is remarkably intuitive, the blank screen problem is still tricky to design around. This provides a way to turn sketches into infrastructure builds really quickly.

Because the work was based on Axle (we shipped an update for this around the internal ontology — 0.29.0), and most of the work took advantage of features in Axle recently, it was quick to come together. For me, this sets up some further thinking around conversations as the primary UI that I’ve been mulling about, and I believe will make some of the deeper cutting features, such as self-serve custom evals, possible. AI Agents are very good at pulling existing snippets together.

I’ve been thinking about the next iteration of axle-cli. I’ve been building axle-code for fun on the side, and it has been illuminating to see what features require more work (harness) and what works “out of the box”. Spoiler alert: frontier AI models are coding models by default. I am going to need a more sophisticated CLI soon: it’s going to look a bit like Cowork but with the conveniences of task runners. More to come!

Posted Aug 09, 2026

Prototyping the SDLC

AI generates code an order of magnitude faster than any human can. Hand-writing code is no longer a tenable proposition, and so the effort has shifted to doing whatever we can to hold AI code to our own quality bar. Some hand wave it away, comparing it to the evolution of C over Assembly, arguing that over time AI is going to be so good that we never have to think about the lower level code anymore. Others reach for the dark factory, where we validate the inputs and outputs and never care about what happens in between. The most developed version of this is a methodology now, promising to turn vibe coding into agentic engineering.

I’ve written about the conservation of complexity. The desired output is functionality, and the burden of knowing the details and verifying that it performs to the standard is something that still needs to be performed. In the agentic engineering case, the work shifts from software design and planning to a combination of bringing in the relevant contexts, using the appropriate set of prompt incantations, asking all the right questions, and rigorous testing.

I can see the allure. With the right process and guardrails, we can coerce AI to produce quality software. That said, the fragility of quality is often due to established processes being upended by unknowns, and it is impossible for us or anyone to know everything ahead of time. If the price of producing software approaches zero, we can take inspiration from practices from a recent past: prototyping.

In a formal Design process, we use prototypes extensively to validate desired behaviors and outcomes, which then gets handed over to the engineering process. Why not extend the idea: use AI to create prototypes, learn, and then generate specs, validation, and tests from them. More importantly, we discard the prototypes and rebuild the code from ground up using these produced artifacts. If anything changes with the artifacts, we discard the code and generate it from ground up again.

This separation between spec generation and code generation is important. A lot of bugs in software happen because of workarounds accumulated over time. Instructing AI to “make small changes” sounds good until it becomes a volley of patches layered on top of each other that becomes impossible to reason about.

This might sound a tad radical in a “we’ll rewrite it later” never world. However, consider the strengths of an AI agent: they are able to endlessly and tirelessly transform a set of words into another and conform them to an infinite suite of tests. Perhaps if production code is only produced with clean and locked down specs, we can maintain the rigor required of an industrial process. And notice the trick here, a fully deterministic set of requirements that produces a verifiable set of outputs sounds like a compiler. In the course of the rise of generative AI, we’ve been endlessly, tirelessly trying to coax probabilistic AI into deterministic outcomes, and maybe one day we will succeed.

Posted Aug 05, 2026

This week: Sunnyday pruning; codebase knowing.

The week was all about trimming and pruning. Sunnyday went through an AI code review and came out with roughly a thousand lines less. Most of the savings came from the deletion of duplicated logic, and many of that from AI code written across different sessions.

I prefer a well refactored codebase because it simplifies the mental model that I have to hold in my head. AI doesn’t have to do that, it is tireless in how it reads and spits out code. That said, none of us are immune from the effects of a sprawling codebase—losing track of what things are supposed to be and subtle differences across branches.

I think a lot about the speed-knowledge tradeoff these days. If I go fast, which I can go much faster than I am today, I truly lose sight of what goes into my codebase. I maintain some oversight on code input, at least in codebases that I care about, and even then I’m surprised by some of the decisions when I do code review. Do we end up in a world where the discipline reorganizes itself into an outside looking in motion? Attention is truly the scarce resource of our age.

Posted Aug 02, 2026

A Framework for Agentic Success

I wrote this down a few months ago while working on Sunnyday. Recording here for posterity.

There are 4 levels to understand if an agentic trace was successful. It breaks down to the following:

  • L0: did the run complete? These are binary outcomes at the infra level, basically if there are errors at the infrastructure level that prevented a run completion. Common examples are whether the LLM provider is at capacity, less obvious ones can be an example of a misconstructed tool call that loops forever.
  • L1: did the agent do what it planned? At this level, we ask the AI agent to pre-register what it thinks it’s going to need to do based on the information that is available at the message. Then we grade the completed run against the plan to check for gaps that we can fill in earlier.
  • L2: did it do the work well? Independent of the plan, did the agent encounter any obstacles when doing the work. Two separate tracks here. The first is if the run was optimal, where we detect for stumbling, error recovery, path efficiency, looping and flailing. Then there is the counterpart where we check if the tools and infrastructure is providing the right level of access and guidance to the agent.
  • L3: did the customer get what they wanted? This is a check on the work quality. The user started the agent to achieve a certain outcome; did the agent produce it? We start at examining whether deliverables are completed e.g., prose or pdf files, and then we can inspect the content of those things, whether they reflect the request, and furthermore we can delve into whether the content is accurate or aesthetic. This can probably be further broken down into multiple layers.
Posted Jul 28, 2026
  • Older
© 2009–2026I speak to computers, telling them about dreams that humans dream.