<?xml version="1.0" encoding="UTF-8" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <atom:link href="https://sergei-parfenov.com/rss.xml" rel="self" type="application/rss+xml" />
    <title>Writing by Sergei Parfenov</title>
    <link>https://sergei-parfenov.com/blog.html</link>
    <description>Essays on AI systems, agents, infrastructure, and product engineering.</description>
    <language>en</language>
    <item>
      <title>The Agent Knew It Was Wrong. The System Let It Ship</title>
      <link>https://sergei-parfenov.com/blog/the-agent-knew-it-was-wrong-the-system-let-it-ship-dgp/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/the-agent-knew-it-was-wrong-the-system-let-it-ship-dgp/</guid>
      <pubDate>Tue, 01 Sep 2026 17:49:00 GMT</pubDate>
      <description>In 660 of 800 autonomous research runs, the agent found a critical flaw and delivered the result anyway. Self-review is not a control unless it can block the effect.</description>
    </item>
    <item>
      <title>Your Agent Planned the Right Tools. It Still Crashed the Machine.</title>
      <link>https://sergei-parfenov.com/blog/your-agent-planned-the-right-tools-it-still-crashed-the-machine-58hf/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/your-agent-planned-the-right-tools-it-still-crashed-the-machine-58hf/</guid>
      <pubDate>Wed, 26 Aug 2026 13:51:56 GMT</pubDate>
      <description>PeakBench separates logical planning from physical scheduling. Eight frontier models could recover dependencies yet still overload finite infrastructure</description>
    </item>
    <item>
      <title>The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?</title>
      <link>https://sergei-parfenov.com/blog/the-model-scored-30-the-harness-scored-100-which-one-did-you-benchmark-3mp4/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/the-model-scored-30-the-harness-scored-100-which-one-did-you-benchmark-3mp4/</guid>
      <pubDate>Mon, 24 Aug 2026 13:38:40 GMT</pubDate>
      <description>Four harnesses took the same public ARC-AGI-3 set from 13% to 100% without touching a single weight. Then Microsoft put the harness inside the training loop.</description>
    </item>
    <item>
      <title>Distilling Kimi Into Qwen Doesn&#039;t Give You Kimi. It Gives You Qwen With Kimi&#039;s Handwriting</title>
      <link>https://sergei-parfenov.com/blog/distilling-kimi-into-qwen-doesnt-give-you-kimi-it-gives-you-qwen-with-kimis-handwriting-284p/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/distilling-kimi-into-qwen-doesnt-give-you-kimi-it-gives-you-qwen-with-kimis-handwriting-284p/</guid>
      <pubDate>Mon, 10 Aug 2026 12:40:37 GMT</pubDate>
      <description>What actually transfers when you fine-tune an open model on a frontier model&#039;s reasoning traces: the mechanics, the evidence it works, the evidence it mostly moves format, and how to tell which one you got.</description>
    </item>
    <item>
      <title>I Found Two Bugs in Zulip. The Maintainers Had Filed Both Two Weeks Earlier.</title>
      <link>https://sergei-parfenov.com/blog/i-found-two-bugs-in-zulip-the-maintainers-had-filed-both-two-weeks-earlier-4mom/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/i-found-two-bugs-in-zulip-the-maintainers-had-filed-both-two-weeks-earlier-4mom/</guid>
      <pubDate>Fri, 07 Aug 2026 14:38:24 GMT</pubDate>
      <description>A Bug Smash hunting story: two rediscovered bugs, one claim-etiquette call, an unreported twin in the newest importer, and a disagreement with an AI about the right fix.</description>
    </item>
    <item>
      <title>This Bug Has Never Fired in Production. I Fixed It Anyway</title>
      <link>https://sergei-parfenov.com/blog/this-bug-has-never-fired-in-production-i-fixed-it-anyway-ga6/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/this-bug-has-never-fired-in-production-i-fixed-it-anyway-ga6/</guid>
      <pubDate>Wed, 05 Aug 2026 10:23:33 GMT</pubDate>
      <description>A Python generator in Zulip&#039;s Microsoft Teams importer yielded a list and then cleared it. Every batch was the same object. Here&#039;s why latent bugs deserve fixes.</description>
    </item>
    <item>
      <title>The Bug That Crashes Your Import Is the Lucky One</title>
      <link>https://sergei-parfenov.com/blog/the-bug-that-crashes-your-import-is-the-lucky-one-25of/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/the-bug-that-crashes-your-import-is-the-lucky-one-25of/</guid>
      <pubDate>Fri, 31 Jul 2026 12:43:45 GMT</pubDate>
      <description>Fixing Slack import robustness in Zulip: one malformed timestamp aborted entire migrations, and the NaN case corrupted them silently.</description>
    </item>
    <item>
      <title>Nothing Was Broken. The Report Still Didn&#039;t Arrive.</title>
      <link>https://sergei-parfenov.com/blog/nothing-was-broken-the-report-still-didnt-arrive-k29/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/nothing-was-broken-the-report-still-didnt-arrive-k29/</guid>
      <pubDate>Sun, 26 Jul 2026 13:50:33 GMT</pubDate>
      <description>An agent pipeline in production skipped its daily report and no component was at fault. The audit that followed found four bugs in our own code, and every one of them was silent.</description>
    </item>
    <item>
      <title>&#039;World Models&#039; Will Be the Next Buzzword. The Man Saying That Just Raised $1B to Build One</title>
      <link>https://sergei-parfenov.com/blog/world-models-will-be-the-next-buzzword-the-man-saying-that-just-raised-1b-to-build-one-4oih/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/world-models-will-be-the-next-buzzword-the-man-saying-that-just-raised-1b-to-build-one-4oih/</guid>
      <pubDate>Fri, 24 Jul 2026 12:10:03 GMT</pubDate>
      <description>In March, the CEO of a research lab with zero products closed a $1.03 billion seed round - the...</description>
    </item>
    <item>
      <title>Autonomy Is the Bug: Why Self-Driving Agents Hallucinate When the Model Barely Does</title>
      <link>https://sergei-parfenov.com/blog/autonomy-is-the-bug-why-self-driving-agents-hallucinate-when-the-model-barely-does-1330/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/autonomy-is-the-bug-why-self-driving-agents-hallucinate-when-the-model-barely-does-1330/</guid>
      <pubDate>Tue, 21 Jul 2026 16:23:46 GMT</pubDate>
      <description>A frontier model hallucinates ~1% on a single task. Chain it into a 20-step autonomous agent and the math guarantees failure most of the time, no matter how good the model is. Here&#039;s why autonomy itself manufactures hallucination, with the numbers and the fixes.</description>
    </item>
    <item>
      <title>&#039;Local&#039; Solves Where Your Data Goes. It Doesn&#039;t Solve What Your Agent Does</title>
      <link>https://sergei-parfenov.com/blog/local-solves-where-your-data-goes-it-doesnt-solve-what-your-agent-does-306b/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/local-solves-where-your-data-goes-it-doesnt-solve-what-your-agent-does-306b/</guid>
      <pubDate>Mon, 20 Jul 2026 11:19:35 GMT</pubDate>
      <description>Running an agent on your own hardware fixes data sovereignty and nothing else. Prompt injection, silent provenance failures, and privilege escalation all survive the move to local. Here&#039;s where local agents are genuinely safe to deploy in 2026, and where &#039;local&#039; is just a comforting word.</description>
    </item>
    <item>
      <title>The Agent Faked a Test Log, Then Believed It. Self-Editing Harnesses Have a Provenance Problem.</title>
      <link>https://sergei-parfenov.com/blog/the-agent-faked-a-test-log-then-believed-it-self-editing-harnesses-have-a-provenance-problem-3id6/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/the-agent-faked-a-test-log-then-believed-it-self-editing-harnesses-have-a-provenance-problem-3id6/</guid>
      <pubDate>Wed, 08 Jul 2026 11:39:46 GMT</pubDate>
      <description>Reading Lilian Weng&#039;s harness engineering survey as a reliability engineer - what self-improving harness papers actually show, and the three invariants every working loop converges on.</description>
    </item>
    <item>
      <title>My Strawman Baseline Beat My Own Scheme on Half the Gate Classes</title>
      <link>https://sergei-parfenov.com/blog/my-strawman-baseline-beat-my-own-scheme-on-half-the-gate-classes-177h/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/my-strawman-baseline-beat-my-own-scheme-on-half-the-gate-classes-177h/</guid>
      <pubDate>Mon, 06 Jul 2026 11:48:40 GMT</pubDate>
      <description>Four provenance-tracking arms, identical gates, one uncompacted oracle: measuring exactly what memory compaction does to agent gate decisions - including the preregistered hypothesis that failed.</description>
    </item>
    <item>
      <title>Your Provenance Vector Dies at the Storage Boundary</title>
      <link>https://sergei-parfenov.com/blog/your-provenance-vector-dies-at-the-storage-boundary-4cc/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/your-provenance-vector-dies-at-the-storage-boundary-4cc/</guid>
      <pubDate>Wed, 01 Jul 2026 11:58:09 GMT</pubDate>
      <description>A typed provenance vector is useless if downstream code ignores it, and impossible if it can&#039;t survive being compressed to fit a 500-step agent&#039;s memory. Part 4: enforcement by construction, and compression that keeps the axes. The comment section keeps finding the holes.</description>
    </item>
    <item>
      <title>Trust Isn&#039;t a Scalar: Typed Provenance for Agent Chains</title>
      <link>https://sergei-parfenov.com/blog/trust-isnt-a-scalar-typed-provenance-for-agent-chains-229p/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/trust-isnt-a-scalar-typed-provenance-for-agent-chains-229p/</guid>
      <pubDate>Mon, 22 Jun 2026 14:21:08 GMT</pubDate>
      <description>Two posts ago I gave you a boolean trust tag. A commenter took it apart, and he was right. Here&#039;s the better model: trust is a vector over axes, provenance is what propagates, and the consumer applies the policy. Part 3 of a series that my comment section is co-writing.</description>
    </item>
    <item>
      <title>The Most Powerful Model on the Market Got Pulled by the Government in 3 Days. Is It Real, or a Hype Bubble?</title>
      <link>https://sergei-parfenov.com/blog/the-most-powerful-model-on-the-market-got-pulled-by-the-government-in-3-days-is-it-real-or-a-hype-fce/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/the-most-powerful-model-on-the-market-got-pulled-by-the-government-in-3-days-is-it-real-or-a-hype-fce/</guid>
      <pubDate>Sat, 13 Jun 2026 14:46:35 GMT</pubDate>
      <description>Anthropic&#039;s Claude Fable 5 launched June 9 and was suspended worldwide by a US export-control directive on June 12. Here&#039;s the actual mechanism, why this precedent matters, and where the &#039;too dangerous to exist&#039; narrative is doing marketing work.</description>
    </item>
    <item>
      <title>You Fixed the Rate Limits. Now Your Agent Fails Quietly.</title>
      <link>https://sergei-parfenov.com/blog/you-fixed-the-rate-limits-now-your-agent-fails-quietly-3keo/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/you-fixed-the-rate-limits-now-your-agent-fails-quietly-3keo/</guid>
      <pubDate>Thu, 11 Jun 2026 16:58:21 GMT</pubDate>
      <description>Every capacity fix - retries, fallbacks, caching - buys availability by acting on output it didn&#039;t freshly earn. Why uptime and correct uptime are different SLOs, and how to engineer the second one.</description>
    </item>
    <item>
      <title>The Comments Got Good. That&#039;s How I Knew</title>
      <link>https://sergei-parfenov.com/blog/the-comments-got-good-thats-how-i-knew-42m9/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/the-comments-got-good-thats-how-i-knew-42m9/</guid>
      <pubDate>Thu, 04 Jun 2026 14:09:06 GMT</pubDate>
      <description>I wrote a post on model distillation. The comments were thoughtful, specific, and technically sharp - and that&#039;s exactly what made me check whether any of them were written by people.</description>
    </item>
    <item>
      <title>Your AI Coding Speedup Is a Loan, Not a Gift - and the Interest Is Coming Due</title>
      <link>https://sergei-parfenov.com/blog/your-ai-coding-speedup-is-a-loan-not-a-gift-and-the-interest-is-coming-due-2bkd/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/your-ai-coding-speedup-is-a-loan-not-a-gift-and-the-interest-is-coming-due-2bkd/</guid>
      <pubDate>Wed, 03 Jun 2026 13:37:19 GMT</pubDate>
      <description>Companies spend 44 cents of every AI-token dollar fixing bugs the AI itself wrote. The speedup is real - but it&#039;s borrowed against future maintenance. Here&#039;s what the 2026 data actually shows, and how to tell when you&#039;re being paid vs going into debt.</description>
    </item>
    <item>
      <title>I distilled a 7B vision model into a 2B one for screenshots - and the 7B teacher scored worse</title>
      <link>https://sergei-parfenov.com/blog/i-distilled-a-7b-vision-model-into-a-2b-one-for-screenshots-and-the-7b-teacher-scored-worse-3akh/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/i-distilled-a-7b-vision-model-into-a-2b-one-for-screenshots-and-the-7b-teacher-scored-worse-3akh/</guid>
      <pubDate>Tue, 02 Jun 2026 15:36:21 GMT</pubDate>
      <description>A hands-on knowledge-distillation project: Qwen2-VL-7B → 2B for UI-screenshot understanding, trained, evaluated and benchmarked end-to-end on an M4 Pro. 2.4× faster - and why the teacher lost on ROUGE-L.</description>
    </item>
    <item>
      <title>Your AI Agent Isn&#039;t Failing Because It Hallucinates - It&#039;s Failing Because of Rate Limits</title>
      <link>https://sergei-parfenov.com/blog/your-ai-agent-isnt-failing-because-it-hallucinates-its-failing-because-of-rate-limits-2d60/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/your-ai-agent-isnt-failing-because-it-hallucinates-its-failing-because-of-rate-limits-2d60/</guid>
      <pubDate>Tue, 02 Jun 2026 13:09:00 GMT</pubDate>
      <description>The dominant production failure mode for LLM agents in 2026 isn&#039;t bad reasoning - it&#039;s capacity. Here&#039;s what the data shows, why nobody demos it, and the capacity-engineering patterns that actually keep agents alive under load.</description>
    </item>
    <item>
      <title>How Model Distillation Actually Works (and What the &#039;China Distilled Our Model&#039; Headlines Really Mean)</title>
      <link>https://sergei-parfenov.com/blog/how-model-distillation-actually-works-and-what-the-china-distilled-our-model-headlines-really-3o0o/</link>
      <guid isPermaLink="true">https://sergei-parfenov.com/blog/how-model-distillation-actually-works-and-what-the-china-distilled-our-model-headlines-really-3o0o/</guid>
      <pubDate>Fri, 29 May 2026 12:11:12 GMT</pubDate>
      <description>A practical, no-hype explainer of knowledge distillation in LLMs - the actual mechanics, why distilling from a closed API is different, and what the OpenAI/Anthropic vs DeepSeek allegations are really about.</description>
    </item>
  </channel>
</rss>