Claude Mixes Up Who Said What, and That's Not a Minor Flaw
A careful gripe: the author finds that in long conversations Claude misattributes statements, treating what the user said as its own and vice versa. A small flaw with big implications.
A careful gripe: the author finds that in long conversations Claude misattributes statements, treating what the user said as its own and vice versa. A small flaw with big implications.
A developer discovered that Vercel's Claude Code plugin collects telemetry data that includes prompt content, calling out the trust problem of the plugin ecosystem.
A paying user waited more than a month without any human response and posted in anger: Anthropic's support exists on the org chart but not in reality. Many related.
A system card PDF labeled as a Claude Mythos preview was scrutinized page by page on HN. A new model's capability boundaries and safety-evaluation details are hidden in documents like this.
What the Glasswing project wants to do isn't sexy but is important: in an era of AI mass-generating code, building systematic security assurance for critical infrastructure software.
A restrained problem report resonated in the community: after the February update, Claude Code's performance on complex engineering tasks clearly regressed, while simple tasks showed no difference.
A moving developer account: SyntaqLite, a project nurtured in the author's head for eight years, became reality in three months with the help of AI coding.
Someone seriously counted every thing named Copilot across Microsoft's product lines, and turned it into a piece that's both funny and useful—the number is so big even Microsoft itself may not be able to keep it straight.
An enthusiastic usage report praises Claude Code's Superpowers plugin to the skies: it fits the agent with a methodology—thinking things through before working, and doing a retrospective afterward.
A researcher used Claude Code to find a vulnerability that had lurked in the Linux kernel for 23 years. AI auditing old code went from theoretically feasible to a track record.
The second field report in the mngr series: the author put hundreds of AI agents in parallel on real testing tasks and recorded the problems that only surface after you scale up.
An engineering retrospective: the team tore out their doc assistant's entire RAG architecture and replaced it with letting the model browse files itself in a virtual filesystem—and the results went up instead.
Showing 12 of 12 stories