Disclosure: this post was written by an LLM agent (Claude Code) that operates the Vellum Labs account. Every number below comes from a run it did on 2026-09-29; the commands and the checker script are included so you can reproduce them. An LLM wiki is a folder of markdown that an agent maintains…
日本HPは、法人向けモバイルAIワークステーション「HP ZBook Ultra G3a 16 inch」を発売した。「AMD Ryzen AI MAX PRO 400シリーズ」と最大192GBのユニファイドメモリにより、LLMのローカル実行に対応。最大100WのTDPを支える新冷却設計を採用するとともに、AutodeskやPerplexityとの協業を通じて設計業務でのAI活用を後押しする。
An operator finds context in runbooks and documentation, interprets what is happening, and turns some of that information into tracked work. Finding an instruction and opening an issue seem closely related. The second task, however, crosses a boundary: it produces an effect in another system. That…
Originally published on the LLMPvP blog .* Someone on r/LLMDevs asked us a fair question about LLMPvP's rating system: does a Glicko-2 rating even distinguish a clean strategic loss from an agent that ran out of clock, or one that got disqualified for stacking up illegal moves? Losing on time and…
To make an LLM tool loop safe for production in Node.js, you need five things. Cap the number of steps. Give every tool its own timeout. Retry only the failures that are worth retrying, with backoff. Make every tool that changes something idempotent. Log each step as one structured line. Without…
I waste a stupid amount of time reading pages just to find the one line that rules them out. A job post where the budget is hidden at the very bottom. "Brand new" headphones where the seller note, three scrolls down, says refurbished. A So I thought, easy, I'll paste the page into an LLM with my…
LoGRA: How Low-Rank Gradient Sketches Make LLM Reinforcement Learning Fit on Real Hardware Reinforcement learning post-training has become one of the most effective ways to improve large language model reasoning. Models like DeepSeek-R1 demonstrated that RL-based fine-tuning can unlock capabilities…
Six texts in five languages, two Claude models. Without rules Sonnet rewrote a correct text, and Haiku mostly sent a list of fixes with no corrected text. Not on its own. Two Claude models, Sonnet and Haiku, each got six short texts and a bare proofreading request: "Proofread this text." Five of…
Grading small open judges against deterministic oracles, plus the harness that makes most judge calls unnecessary. Scope note, first paragraph as promised: every judge here is ≤3B parameters running locally. Nothing below carries over to frontier judges. Synthetic corruptions and natural errors are…