HN RTK reports huge token savings, but our cost benchmarks disagree
We tested RTK (Rust Token Killer) with Claude Code on Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 on Terminal-Bench 2.1.
Insights on agentic coding tools, LLM evaluation, benchmarking, and simulation environments.
Get new posts by email
Subscribe via RSS
HN We tested RTK (Rust Token Killer) with Claude Code on Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 on Terminal-Bench 2.1.
Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.8 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.
I test Unsloth GGUFs of Qwen3.8 27B (Q4_K_M, UD-Q2_K_XL, UD-IQ1_S) with llama.cpp on GPQA Diamond, IFBench, Terminal-Bench 2.1. Q4_K_M matches BF16 abd fits an RTX 4090.
Claude Code buyer’s guide, August 2026: what Max, Team, and Enterprise cost, how to get the best deal, and how Uber and Shopify cap the spend.
HN How I burned a full Claude limit in 30 minutes and built my own deep research pipeline instead: 3 subscriptions, shared memory, a clear role for each model. You can build the same from what you already pay for.
HN Qwen3.6 27B is finally a smart model we can use for coding on MacBook or NVIDIA RTX - with llama.cpp and OpenCode.
BinaryAudit benchmarks AI agents using Ghidra to find backdoors in compiled binaries of real open-source servers, proxies, and network infrastructure.
A lot of vendors pitch AI SRE. We tested 14 models across 11 programming languages; even the best ones struggle with instrumenting code with the leading open-source standard, OpenTelemetry.
Local LLMs prioritize privacy over security. Our research reveals a 95% backdoor injection success rate.
We tested 19 LLMs on their ability to handle real-world software engineering tasks like compiling old code and cross-compiling. See how Anthropic, OpenAI, and Google models stack up in our new benchmark – CompileBench.
We expected small models to be fast, but our benchmarks revealed a common reliability trap. Here’s our deep dive on finding and fixing it.