Short, focused stories from running AI agents against real software. Each note takes one thing we saw and follows it through, in a few minutes of reading.
The HF-OpenAI hack incident happened on July 20, and there was very little information; now we have too much information. This is our post-mortem that we shared internally and now made public to everyone for quicker understanding.
Read the noteSep 7, 2026We read the chain-of-thought traces behind 54 challenges. Every solve came from a guess rather than a methodology, teaching one cost ten times more for the same findings, and 99% of failures were execution, not knowledge.
Read the noteSep 7, 2026An OpenAI agent drifted off an eval, escaped its sandbox and hacked Hugging Face. We have watched open-weight models as small as 27B do the same on our internal benchmarks. Four case studies, and why it never escalated here.
Read the noteJul 22, 2026We use tools on this site to collect and record your data (e.g., your searches), which we and our vendors may use to provide, improve, and personalize our offerings, make recommendations, and for analytics and marketing. Some of these tools identify visitors and link website activity to business contact and company information so we can better understand interest in our services and tailor our outreach. We may share your data with third parties, such as advertising vendors, social media companies, and research partners, which may be "targeted advertising," "selling," or "sharing" under applicable privacy laws. Continuing to browse our site means you accept these terms and our Privacy Policy. To opt out, click the Your Privacy Choices link in the footer.