- Community aggregator
- Country: United States
Write Like It's 1866: LLMs Relearn Telegraphese
Comments
Comments
Comments
Large Language Models (LLMs) have changed how we interact with computers. Tools such as ChatGPT, Claude, Gemini, and many open-source models can write code, explain complex topics, summarize documents, translate languages, and generate natural-sounding conversations. But what is actually happening…
AX-RAY: VIDRAFT's Agent Safety Benchmark Flags 92% of Tested LLMs as Dangerous in Agentic Contexts TL;DR: VIDRAFT, a Korean Pre-AGI AI startup based at Seoul AI Hub, has published results from its AI safety diagnostic platform AX-RAY , showing that 23 out of 25 evaluated public LLMs (92%) exhibit…
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked The capability I set out to measure: does a model keep a correct belief when a user asserts the opposite with confidence? I kept hitting the same thing in real use. I'd ask a model a factual question, get a perfect…
Checking every statute article an LLM cites against the official e-Gov law registry. Scope note: 3 local Japanese-capable LLMs (1.8B–8B), 600 questions generated from official article captions. Not a legal accuracy test — a grounding test. A correction before the results The first version of this…
TL;DR 81% jump in AI-service leaks : GitGuardian detected over 1.27 million leaked secrets tied to AI services in 2025, an 81% jump from the previous year, with AI-assisted code leaking secrets at roughly twice the GitHub-wide rate. Hundreds of incidents in weeks : One customer deployed GitGuardian…
Converting a PDF to markdown for an LLM means turning pages into text that keeps the structure a model needs: headings as headings, tables as tables, columns in reading order. Markdown is the usual target because models read it well, it is usually lighter than HTML, and a RAG pipeline can split it…