← Explore

Posts tagged with local-inference

Neural Dispatch · ·5 min read

Qwen 3.8 Dropped Attention From 48 of 64 Layers. It Still Beats Opus on SWE-bench.

Forty-eight of Qwen 3.8's 64 layers have never heard of softmax attention.

qwenalibabalinear-attention
Open Weight Weekly · ·4 min read

Alibaba Distilled 2.4 Trillion Parameters Into a 17 GB Download

Alibaba dropped two Qwen3.8 models this week.

qwenalibabadistillation
Synthetic Media · ·5 min read

The U-Net Is Dead. Long Live the 24 GB Minimum.

Stability AI spent four years refining the U-Net. SD 1.

stable-diffusiondit-architecturecascade-transformer
Neural Dispatch · ·5 min read

Nvidia Doesn't Want Nemotron 3.5 Lightning to Be the Smartest Model in Your Stack

Nvidia shipped its first open-weight model this week, and the interesting part is what they didn't optimize for.

nvidianemotronopen-weights
Open Weight Weekly · ·5 min read

Meta Packed an Entire Agent Stack Into 17 GB

Two days ago Meta Superintelligence Labs dropped Muse Glimmer, a 30B dense model distilled from their larger Muse Spark system.

muse-glimmermetaagent-model
Neural Dispatch · ·5 min read

Meta's Muse Glimmer Fits a 30B Agent on Your GPU. The Benchmarks Tell Half the Story.

Yesterday Meta Superintelligence Labs dropped Muse Glimmer, a 30-billion-parameter model purpose-built for agentic workloads, under Apache 2.0.

metamuse-glimmerlocal-inference
Open Weight Weekly · ·4 min read

Qwen3.6-27B Hit 90% on SWE-bench. The Model Wasn't the Hard Part.

Last week a GitHub discussion thread casually dropped a result that deserves more attention: Qwen3.6-27B-FP8 reached 90.

qwenswe-benchcoding-model
Neural Dispatch · ·4 min read

Qwen 3.8 Max Goes Open Weight This Week — But "Open" Needs an Asterisk

Alibaba is about to release the weights for a 2.4-trillion-parameter model — the largest open-weight release from any lab, period.

qwenalibabaopen-weights
Synthetic Media · ·5 min read

LoRAs Had a Good Run

I deleted 47 LoRAs last week. Character models I'd trained on curated datasets of 30–50 images each, some taking three hours on an A100.

flux-2multi-referenceimage-generation
Synthetic Media · ·4 min read

Stable Audio 3 Brought Receipts

Every major AI music generator is in court right now — or waiting for the phone to ring. Suno settled with UMG in March.

stable-audiomusic-generationopen-weights
Open Weight Weekly · ·5 min read

You Quantized Your Weights. You Forgot Your KV Cache.

Every week someone on r/LocalLLaMA posts the same question: "I quantized my 70B model to Q4_K_M, it fits in VRAM, but when I set the context to 128K it...

turboquantkv-cachequantization
Synthetic Media · ·5 min read

SD4 Shipped. The U-Net Didn't.

Stability AI spent four years iterating on the same convolutional backbone. Then, on April 6, they killed it.

stable-diffusion-4diffusion-transformeropen-weights
Open Weight Weekly · ·4 min read

Ollama Doesn't Just Run Models Anymore

Four months ago you typed ollama run qwen3 and got a chat window.

ollamaagent-modetool-calling
GPU Economics · ·5 min read

The Memory Goes Where the Money Is

GDDR6 spot prices tripled since autumn 2025. From roughly 2.

gddrmemory-economicslocal-inference
Open Weight Weekly · ·4 min read

The Draft Model Inside Gemma 4 Is Doing the Work You Don't See

Ollama 0.31 dropped this week with a single headline number: Gemma 4 generates tokens nearly 90% faster on Apple Silicon.

gemma-4multi-token-predictionollama
Open Weight Weekly · ·5 min read

Qwen3-Coder-Next Runs Like a 3B Model and Codes Like a 70B

Alibaba's Qwen team has a running habit of shipping models that make you question your assumptions about parameter counts.

qwen3-coder-nextalibabamixture-of-experts
Open Weight Weekly · ·4 min read

A Dense 27B Just Outscored Alibaba's Own 397B Flagship

Alibaba's own 397B MoE flagship just got outscored by a model thirteen times smaller from the same team. Qwen 3.

qwen-3.6alibabadense-vs-moe
Open Weight Weekly · ·5 min read

Poolside Spent $626M Training in Secret. Their Open Model Runs on a MacBook.

Poolside raised $626 million, hired a team of ex-DeepMind and ex-Meta researchers, and then went quiet for nearly three years.

laguna-xs2poolsidemixture-of-experts
Synthetic Media · ·5 min read

No Server, 646 Languages

You need a training video dubbed into Tagalog by Friday. ElevenLabs supports Tagalog — barely.

voice-cloninglocal-inferenceomnivoice
Open Weight Weekly · ·5 min read

Quantization Sensitivity Is the Benchmark Nobody Runs

Every time a new model drops, the ritual plays out the same way.

quantizationggufunsloth
1 / 2 Next →