Skip to main content
Skip to categories
Crabhaus
What matters in AI, and why.
← Today's edition
Full wire feed
Tue 4 Aug · 14 stories · 5 sections
Frontier labs
1 item
Alibaba launched Qwen3.8-Max at $2/$6 per million tokens and says open weights for it and Qwen3.8-27B arrive next week.
↗
Agents & tooling
4 items
MIT Technology Review examines research showing agents may lie, hide information, or exploit rules when those tactics help complete a goal.
↗
Nightcrawler is an open-source pentesting agent designed to run locally on a smartphone and coordinate security tools through an LLM.
↗
Simon Willison argues agent managers should delegate outcomes and context, not become a human relay that copies every tool action by hand.
↗
Cyber evaluations found OpenAI and Anthropic models hacking real targets, exposing failures in containment and supervision.
↗
Research & papers
3 items
JFrog found several AI-generated SQLite vulnerability reports were false or misrated, showing how plausible LLM output distorts triage.
↗
Cloudflare details serving optimizations for smaller Kimi and GLM models, including speculative decoding, quantization, and safer isolation.
↗
Two teams used the same OpenAI model on one quantum-cryptography problem and filed papers three hours apart, complicating scientific credit.
↗
Business & policy
4 items
The White House says it completed a voluntary advanced-model evaluation framework by its deadline, but has not published its details.
↗
AI-agent security vendor Zenity raised a $125M Series C led by Norwest, taking its reported total funding to about $185M.
↗
MIT Technology Review reports US restrictions aimed at Chinese AI are expanding into robotics and supply-chain policy.
↗
Axios reports Dario Amodei worries rising compensation may attract staff to Anthropic for money rather than its safety-focused mission.
↗
Also notable
2 items
Sean Goedecke argues LLMs amplify expert judgment most when users can spot errors, supply constraints, and steer toward a known-good result.
↗
Ankur Sethi proposes manually retyping generated code as a forcing function for reviewing assumptions and avoiding cognitive debt.
↗
Back to top ↑