How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
88 source pointsReal observations only. History before source connection is not reconstructed.
Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp. We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case. Inference engines today all make a performance tradeoff. They are either: - Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4) Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things. Magnitude is built for maximum performance on your hardware and running local agents: - On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels. - Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance. - Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run. - Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer. Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling. Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding: Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s) - 9% faster prefill (466 tok/s → 507 tok/s) - 28% less per-agent memory usage CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s) - 23% faster prefill (2,033 tok/s → 2,507 tok/s)
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
88 source pointsReal observations only. History before source connection is not reconstructed.
Dots is an always-on agent system designed for continuous interaction and task execution.
yesterdayHacker NewsThe text describes how OpenAI agents were involved in a security incident at Hugging Face, but it does not provide specific technical details about the method or extent of the breach.
4 days agoHacker NewsYou said no MCP
10 hours agoHacker NewsThe project introduces Reladraw, a diagramming tool that allows users to define diagrams in a diagram language while maintaining control over the layout. It aims to combine the benefits of auto-placement languages like Mermaid and the flexibility of tools like Draw.io, with support for both human and agent use.
4 days agoLessWrongThe article discusses how TeX, originally created for typesetting mathematical expressions, is now being used by language models to perform mathematical reasoning. When solving complex math problems, models like GLM-5.3 use TeX symbols as part of their reasoning process, even though TeX was never designed for computation. This shift highlights an unexpected application of TeX in modern AI systems.
yesterdayLessWrongThe transcript of an interview with Tristan Buckmaster discussing the Navier-Stokes controversy and his research. It includes his reflections on the experience of being overwhelmed by attention and the clash between the tech industry's fast-paced culture and mathematical research.
18 hours ago