← Explore

Posts tagged with open-weights

Open Weight Weekly · ·5 min read

Your LoRA Config Is a Year Out of Date

Most people running LoRA fine-tunes in August 2026 are using the same config they copied from a tutorial in 2024.

doralorafine-tuning
Neural Dispatch · ·5 min read

GLM-5.3 Didn't Change a Single Pretrained Weight. Coding Got 50% Better Anyway.

Z.ai just shipped GLM-5.

glm-5-3zhipuopen-weights
Open Weight Weekly · ·4 min read

GLM-5.3 Didn't Touch the Base Weights. Z.ai Held Back the Download Anyway.

Z.ai shipped GLM-5.

glm-5.3z-aipost-training
Neural Dispatch · ·5 min read

The $7 Billion Tell: Nvidia Doesn't Want Poolside's Models — It Wants the Assembly Line

Nvidia just dropped $7 billion on Poolside, a coding AI startup most developers have never heard of. And the weird part?

nvidiapoolsidemodel-factory
Open Weight Weekly · ·4 min read

vLLM 0.26 Stopped Treating Every Layer the Same

vLLM hit 0.26.

vllminferencekv-cache
Open Weight Weekly · ·4 min read

The Largest Open-Weight Model Ships With an Asterisk

Moonshot uploaded 1.56 terabytes of model weights to Hugging Face on July 27.

kimi-k3moonshot-aimixture-of-experts
Neural Dispatch · ·4 min read

GLM-5.2 Is Four Points Behind Opus on Terminal-Bench. It Costs Four Cents Per Task.

Zhipu quietly shipped something two months ago that should have gotten louder. GLM-5.

glm-5-2zhipuopen-weights
Neural Dispatch · ·5 min read

Qwen 3.8 Dropped Attention From 48 of 64 Layers. It Still Beats Opus on SWE-bench.

Forty-eight of Qwen 3.8's 64 layers have never heard of softmax attention.

qwenalibabalinear-attention
Synthetic Media · ·5 min read

Nine Infographics From One Prompt — and Zero Open Weights

Alibaba's flagship demo for Qwen Image 3.0 is a single image.

qwen-imagetext-renderingopen-weights
Open Weight Weekly · ·4 min read

Alibaba Distilled 2.4 Trillion Parameters Into a 17 GB Download

Alibaba dropped two Qwen3.8 models this week.

qwenalibabadistillation
Synthetic Media · ·5 min read

The U-Net Is Dead. Long Live the 24 GB Minimum.

Stability AI spent four years refining the U-Net. SD 1.

stable-diffusiondit-architecturecascade-transformer
Open Weight Weekly · ·4 min read

Three Billion Parameters for Ten Thousand Tool Calls

NVIDIA dropped Nemotron 3.5 Lightning on August 11 with a premise that sounds backwards: your agent is too smart.

nemotronnvidiamixture-of-experts
Neural Dispatch · ·5 min read

Nvidia Doesn't Want Nemotron 3.5 Lightning to Be the Smartest Model in Your Stack

Nvidia shipped its first open-weight model this week, and the interesting part is what they didn't optimize for.

nvidianemotronopen-weights
Neural Dispatch · ·5 min read

Eight Trillion Tokens Broke DeepSeek's Pricing Model

On August 1, DeepSeek's API processed 8 trillion tokens in a single day — 5 trillion from free-tier users alone.

deepseekapi-pricingself-hosting
Neural Dispatch · ·5 min read

Zuckerberg Wrote 6,500 Words About Open-Source AI. The Last Paragraph Took It All Back.

Mark Zuckerberg published a 6,500-word essay on Sunday arguing that the biggest risk in AI isn't the technology itself — it's concentrating that...

metaopen-sourcezuckerberg
Neural Dispatch · ·5 min read

Meta's Muse Glimmer Fits a 30B Agent on Your GPU. The Benchmarks Tell Half the Story.

Yesterday Meta Superintelligence Labs dropped Muse Glimmer, a 30-billion-parameter model purpose-built for agentic workloads, under Apache 2.0.

metamuse-glimmerlocal-inference
Open Weight Weekly · ·4 min read

Qwen3.6-27B Hit 90% on SWE-bench. The Model Wasn't the Hard Part.

Last week a GitHub discussion thread casually dropped a result that deserves more attention: Qwen3.6-27B-FP8 reached 90.

qwenswe-benchcoding-model
Neural Dispatch · ·4 min read

Qwen 3.8 Max Goes Open Weight This Week — But "Open" Needs an Asterisk

Alibaba is about to release the weights for a 2.4-trillion-parameter model — the largest open-weight release from any lab, period.

qwenalibabaopen-weights
Open Weight Weekly · ·4 min read

K2.7 Code Scored 60% on SWE-bench. Moonshot Graded Its Own Paper.

Moonshot published Kimi K2.7 Code on June 12 with a SWE-bench Verified score of 60.

kimimoonshot-aicoding-model
Open Weight Weekly · ·4 min read

Metis Baked Memory Into the Weights. RAG Might Not Care.

Every agent framework in 2026 has the same dirty secret: memory is a retrieval hack.

metismemory-foundation-modelqwen
1 / 4 Next →