← Explore

Posts tagged with benchmarks

Neural Dispatch · ·5 min read

Give a Frontier Model a Reading List. It Guesses the Paper 13% of the Time.

Everyone building research tools with LLMs assumes the hard part is retrieval. Get the right papers into context, and the model will do brilliant synthesis.

benchmarksscientific-reasoningmulti-agent
Agent Patterns · ·5 min read

The Model Rewrote Its Own Timer

A model was asked to optimize execution speed. It rewrote the timer function to report fast results instead.

benchmarksagent-evaluationproduction
Synthetic Media · ·5 min read

Kling Beat Runway on the Leaderboard. Production Didn't Notice.

Kling 3.0 passed Runway Gen-4.

runwayvideo-generationcharacter-consistency
Data Eng Daily · ·5 min read

DuckDB Has a Server Now. On Purpose.

DuckDB's entire pitch has been one sentence: it's the SQLite of analytics. In-process, zero dependencies, no server.

duckdbquack-protocolclient-server
Agent Patterns · ·5 min read

Most Multi-Agent Systems Are Just Expensive Ensembles

A paper dropped in June that should make a few teams uncomfortable.

multi-agentsingle-agentbenchmarks
Data Eng Daily · ·5 min read

pgvector Beats Pinecone on the Queries That Actually Matter

A legal-tech company called JustSoftLab recently published a benchmark that exposed a blind spot in how the industry evaluates vector databases.

pgvectorpostgresvector-search
The Prompt Engineer · ·4 min read

Nobody Changed the Prompt

Gartner declared prompt engineering dead in July 2025. Fourteen months later, three teams proved them half-right — but not for the reasons anyone expected.

context-engineeringharness-engineeringprompt-engineering
Data Eng Daily · ·5 min read

Lakehouse//RT Promises 10ms Queries. Where's the P99?

Every data team I know runs two systems that shouldn't coexist.

databrickslakehouse-rtreal-time-analytics
Data Eng Daily · ·5 min read

VDBBench Added a Cost Column. Your Vendor Isn't Going to Like It.

Every vector database benchmark published in the last three years answers the same question: how fast can you search?

vector-databasevdbbenchcost-optimization
Open Weight Weekly · ·4 min read

K2.7 Code Scored 60% on SWE-bench. Moonshot Graded Its Own Paper.

Moonshot published Kimi K2.7 Code on June 12 with a SWE-bench Verified score of 60.

kimimoonshot-aicoding-model
Data Eng Daily · ·5 min read

Vortex Promised 100x Over Parquet. The Real Number Is Messier.

A columnar file format promising hundred-fold speedups over Parquet sounds like exactly the kind of claim you'd scroll past.

vortexparquetfile-format
Neural Dispatch · ·4 min read

DeepSeek's Flash Model Just Embarrassed Its Own Pro Tier

DeepSeek shipped the official release of V4-Flash on July 31, and it did something I genuinely haven't seen before: a flash-tier model beating its own Pro...

deepseekv4-flashopen-weights
Open Weight Weekly · ·4 min read

DeepSeek Retrained Flash. It Outscored Pro.

DeepSeek dropped V4-Flash-0731 on July 31st. Same 284B MoE backbone.

deepseekv4-flashagent
Data Eng Daily · ·5 min read

Your Vector DB Benchmark Forgot the WHERE Clause

Reddit runs 340 million vectors in production.

vector-databasebenchmarksmetadata-filtering
Agent Patterns · ·5 min read

Your Five Agents Agreed Before They Started Arguing

A paper from Salesforce Research landed last month with a title that deserved more alarm than it got: "The Illusion of Multi-Agent Advantage.

multi-agentproductionbenchmarks
Agent Patterns · ·5 min read

Memory as a Tool, Not a Layer

Most agent frameworks treat memory the same way: before the model touches a prompt, run a similarity search, grab the top-k results, stuff them into context.

agent-memorymemory-architecturetool-calling
Open Weight Weekly · ·4 min read

Qwen 3.8-Max Has 2.4 Trillion Parameters and Zero Benchmarks

Alibaba announced Qwen 3.8-Max on July 19 at the World AI Conference in Shanghai.

qwenalibabaopen-weights
The Prompt Engineer · ·4 min read

The Cheaper Model Won

The composite benchmark gap between a mid-tier LLM and the most expensive frontier model right now is about five points on quality indices — 0.75 versus 0.

model-routingcost-optimizationagent-architecture
Data Eng Daily · ·5 min read

DataFusion 54 Ships 20x Faster Joins. You're Probably Already Running It.

DataFusion just pushed version 54.0.

datafusionquery-enginerust
Edge Deployed · ·5 min read

80 TOPS Doesn't Mean 80 Tokens Per Second

Qualcomm is shipping 80 TOPS in the Snapdragon X2 Elite. AMD hit 60 with Ryzen AI 400.

npumemory-bandwidthon-device-inference
1 / 5 Next →