# Sergei Parfenov > Sergei Parfenov, CTO at Aliwio and Symptomato. Early Yandex Praktikum team member building AI products, learning infrastructure and automated project assessment. This is the personal site and article archive of Sergei Parfenov. All text is public and available without JavaScript or authentication. Article metadata includes the author, publication date, canonical URL, and DEV counterpart when present. ## Profile - [Profile](https://sergei-parfenov.com/index.md): Biography, current roles, full career history, interests, and contact links. - [Profile JSON](https://sergei-parfenov.com/profile.json): Structured public profile. ## Articles - [Article index](https://sergei-parfenov.com/articles.json): Every published article with its canonical URL and Markdown URL. - [Blog](https://sergei-parfenov.com/blog.html): Human-readable archive. - [RSS](https://sergei-parfenov.com/rss.xml): Recent publications. - [The Agent Knew It Was Wrong. The System Let It Ship](https://sergei-parfenov.com/blog/the-agent-knew-it-was-wrong-the-system-let-it-ship-dgp/index.md): In 660 of 800 autonomous research runs, the agent found a critical flaw and delivered the result anyway. Self-review is not a control unless it can block the effect. - [Your Agent Planned the Right Tools. It Still Crashed the Machine.](https://sergei-parfenov.com/blog/your-agent-planned-the-right-tools-it-still-crashed-the-machine-58hf/index.md): PeakBench separates logical planning from physical scheduling. Eight frontier models could recover dependencies yet still overload finite infrastructure - [The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?](https://sergei-parfenov.com/blog/the-model-scored-30-the-harness-scored-100-which-one-did-you-benchmark-3mp4/index.md): Four harnesses took the same public ARC-AGI-3 set from 13% to 100% without touching a single weight. Then Microsoft put the harness inside the training loop. - [Distilling Kimi Into Qwen Doesn't Give You Kimi. It Gives You Qwen With Kimi's Handwriting](https://sergei-parfenov.com/blog/distilling-kimi-into-qwen-doesnt-give-you-kimi-it-gives-you-qwen-with-kimis-handwriting-284p/index.md): What actually transfers when you fine-tune an open model on a frontier model's reasoning traces: the mechanics, the evidence it works, the evidence it mostly moves format, and how to tell which one you got. - [I Found Two Bugs in Zulip. The Maintainers Had Filed Both Two Weeks Earlier.](https://sergei-parfenov.com/blog/i-found-two-bugs-in-zulip-the-maintainers-had-filed-both-two-weeks-earlier-4mom/index.md): A Bug Smash hunting story: two rediscovered bugs, one claim-etiquette call, an unreported twin in the newest importer, and a disagreement with an AI about the right fix. - [This Bug Has Never Fired in Production. I Fixed It Anyway](https://sergei-parfenov.com/blog/this-bug-has-never-fired-in-production-i-fixed-it-anyway-ga6/index.md): A Python generator in Zulip's Microsoft Teams importer yielded a list and then cleared it. Every batch was the same object. Here's why latent bugs deserve fixes. - [The Bug That Crashes Your Import Is the Lucky One](https://sergei-parfenov.com/blog/the-bug-that-crashes-your-import-is-the-lucky-one-25of/index.md): Fixing Slack import robustness in Zulip: one malformed timestamp aborted entire migrations, and the NaN case corrupted them silently. - [Nothing Was Broken. The Report Still Didn't Arrive.](https://sergei-parfenov.com/blog/nothing-was-broken-the-report-still-didnt-arrive-k29/index.md): An agent pipeline in production skipped its daily report and no component was at fault. The audit that followed found four bugs in our own code, and every one of them was silent. - ['World Models' Will Be the Next Buzzword. The Man Saying That Just Raised $1B to Build One](https://sergei-parfenov.com/blog/world-models-will-be-the-next-buzzword-the-man-saying-that-just-raised-1b-to-build-one-4oih/index.md): In March, the CEO of a research lab with zero products closed a $1.03 billion seed round - the... - [Autonomy Is the Bug: Why Self-Driving Agents Hallucinate When the Model Barely Does](https://sergei-parfenov.com/blog/autonomy-is-the-bug-why-self-driving-agents-hallucinate-when-the-model-barely-does-1330/index.md): A frontier model hallucinates ~1% on a single task. Chain it into a 20-step autonomous agent and the math guarantees failure most of the time, no matter how good the model is. Here's why autonomy itself manufactures hallucination, with the numbers and the fixes. - ['Local' Solves Where Your Data Goes. It Doesn't Solve What Your Agent Does](https://sergei-parfenov.com/blog/local-solves-where-your-data-goes-it-doesnt-solve-what-your-agent-does-306b/index.md): Running an agent on your own hardware fixes data sovereignty and nothing else. Prompt injection, silent provenance failures, and privilege escalation all survive the move to local. Here's where local agents are genuinely safe to deploy in 2026, and where 'local' is just a comforting word. - [The Agent Faked a Test Log, Then Believed It. Self-Editing Harnesses Have a Provenance Problem.](https://sergei-parfenov.com/blog/the-agent-faked-a-test-log-then-believed-it-self-editing-harnesses-have-a-provenance-problem-3id6/index.md): Reading Lilian Weng's harness engineering survey as a reliability engineer - what self-improving harness papers actually show, and the three invariants every working loop converges on. - [My Strawman Baseline Beat My Own Scheme on Half the Gate Classes](https://sergei-parfenov.com/blog/my-strawman-baseline-beat-my-own-scheme-on-half-the-gate-classes-177h/index.md): Four provenance-tracking arms, identical gates, one uncompacted oracle: measuring exactly what memory compaction does to agent gate decisions - including the preregistered hypothesis that failed. - [Your Provenance Vector Dies at the Storage Boundary](https://sergei-parfenov.com/blog/your-provenance-vector-dies-at-the-storage-boundary-4cc/index.md): A typed provenance vector is useless if downstream code ignores it, and impossible if it can't survive being compressed to fit a 500-step agent's memory. Part 4: enforcement by construction, and compression that keeps the axes. The comment section keeps finding the holes. - [Trust Isn't a Scalar: Typed Provenance for Agent Chains](https://sergei-parfenov.com/blog/trust-isnt-a-scalar-typed-provenance-for-agent-chains-229p/index.md): Two posts ago I gave you a boolean trust tag. A commenter took it apart, and he was right. Here's the better model: trust is a vector over axes, provenance is what propagates, and the consumer applies the policy. Part 3 of a series that my comment section is co-writing. - [The Most Powerful Model on the Market Got Pulled by the Government in 3 Days. Is It Real, or a Hype Bubble?](https://sergei-parfenov.com/blog/the-most-powerful-model-on-the-market-got-pulled-by-the-government-in-3-days-is-it-real-or-a-hype-fce/index.md): Anthropic's Claude Fable 5 launched June 9 and was suspended worldwide by a US export-control directive on June 12. Here's the actual mechanism, why this precedent matters, and where the 'too dangerous to exist' narrative is doing marketing work. - [You Fixed the Rate Limits. Now Your Agent Fails Quietly.](https://sergei-parfenov.com/blog/you-fixed-the-rate-limits-now-your-agent-fails-quietly-3keo/index.md): Every capacity fix - retries, fallbacks, caching - buys availability by acting on output it didn't freshly earn. Why uptime and correct uptime are different SLOs, and how to engineer the second one. - [The Comments Got Good. That's How I Knew](https://sergei-parfenov.com/blog/the-comments-got-good-thats-how-i-knew-42m9/index.md): I wrote a post on model distillation. The comments were thoughtful, specific, and technically sharp - and that's exactly what made me check whether any of them were written by people. - [Your AI Coding Speedup Is a Loan, Not a Gift - and the Interest Is Coming Due](https://sergei-parfenov.com/blog/your-ai-coding-speedup-is-a-loan-not-a-gift-and-the-interest-is-coming-due-2bkd/index.md): Companies spend 44 cents of every AI-token dollar fixing bugs the AI itself wrote. The speedup is real - but it's borrowed against future maintenance. Here's what the 2026 data actually shows, and how to tell when you're being paid vs going into debt. - [I distilled a 7B vision model into a 2B one for screenshots - and the 7B teacher scored worse](https://sergei-parfenov.com/blog/i-distilled-a-7b-vision-model-into-a-2b-one-for-screenshots-and-the-7b-teacher-scored-worse-3akh/index.md): A hands-on knowledge-distillation project: Qwen2-VL-7B → 2B for UI-screenshot understanding, trained, evaluated and benchmarked end-to-end on an M4 Pro. 2.4× faster - and why the teacher lost on ROUGE-L. - [Your AI Agent Isn't Failing Because It Hallucinates - It's Failing Because of Rate Limits](https://sergei-parfenov.com/blog/your-ai-agent-isnt-failing-because-it-hallucinates-its-failing-because-of-rate-limits-2d60/index.md): The dominant production failure mode for LLM agents in 2026 isn't bad reasoning - it's capacity. Here's what the data shows, why nobody demos it, and the capacity-engineering patterns that actually keep agents alive under load. - [How Model Distillation Actually Works (and What the 'China Distilled Our Model' Headlines Really Mean)](https://sergei-parfenov.com/blog/how-model-distillation-actually-works-and-what-the-china-distilled-our-model-headlines-really-3o0o/index.md): A practical, no-hype explainer of knowledge distillation in LLMs - the actual mechanics, why distilling from a closed API is different, and what the OpenAI/Anthropic vs DeepSeek allegations are really about. ## MCP - [Connection details](https://sergei-parfenov.com/mcp.json): Public read-only Streamable HTTP endpoint at https://sergei-parfenov.com/mcp. Tools: get_profile, search, fetch. No API key is needed. ## Source and citation Use each document's canonical HTML URL for a link or citation. The Markdown and JSON URLs are alternate representations, not separate publications. Drafts and the retired resume PDF are not exposed. This file is a convenience index for agents, not a Google ranking signal.