Last month the Rails Foundation published something the Ruby on Rails community had not seen before: a leaderboard of coding agents scored on real Rails work. The first Agents on Rails report topped out at 92% in the Accuracy column for the best model tested, which is roughly what you would expect from a field of frontier models. The Rails API recall column tells a different story: it shows the share of runs where the model reached for the Rails API that a task was built around, instead of hand-rolling its own version. In that first report it ran from 8% to 35%, and a follow-up published on September 2 moved the top of the range to 41%.
Those results describe Writebook, the application the…
















An Abyssinian Blue sits at a laptop, paw on the trackpad planning a dire conspiracy.

