Dots are designed to automate online tasks, like buying furniture. In my initial experience, the always-on agent was a bit buggy and couldn’t complete a captcha.
Wikimedia Foundation just published findings from their investigation into OpenAI agent activity across their platforms. The forensic trail includes sandbox edits, Etherpad exploitation attempts, and hundreds of thousands of unauthorized queries to the Wikidata Query Service. The activity started…
A custom AI lead qualification agent costs between 15k and 60k USD to build and somewhere from a few hundred to a few thousand dollars a month to run. The wide range is not about the model. It comes from how many data sources the agent has to read, how deep the CRM integration goes, and how much of…
A few months back, I was setting up a Central Monitoring system for all my projects where I could do log monitoring, cost monitoring, compute utilisation by each project, and so on. For that, I had already set up an ML-based algorithm to detect anomalies in cost and RAM, which then sent me alerts…
I spend a lot of time in the AI space -- reading papers, building things, talking to engineers who are actually shipping. And there is a gap between what the demos show and what production systems actually look like that nobody is being fully honest about. So here is my honest take on where things…
The 3 AM page Picture this. You get paged at 3 AM for a production outage. A teammate hands you one log file with the exact error in it. You find the problem in five minutes. Now replay the same night. This time your teammate hands you that same log file, plus 500 unrelated log files, three months…
There's a line from Hacker News this week that reads like a shrug but is actually a thesis: "We're going to need default hard budget caps on pretty much everything." (76 pts) Read it again. Not "on AI." On pretty much everything. The person isn't talking about a feature request. They're describing…
OpenAI and Ironclad just published a case study on using production contract workflows as both training data and evaluation benchmarks for computer-use agents. This is not a demo. It is a partnership where a SaaS company opens its workflow engine to become an agent training ground. The plumbing…
A user abandoned a research report at second 41. The pipeline was still running. Five agents, each waiting for the previous one to finish, even though three of them had no actual dependency on each other. By the time the output landed, the tab was closed. Sequential orchestration looks reasonable…