<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>Measured Locally</title>
<subtitle>Local AI, timed on machines I own</subtitle>
<link href="https://measuredlocally.com/"/>
<link rel="self" href="https://measuredlocally.com/feed.xml"/>
<id>https://measuredlocally.com/</id>
<updated>2026-10-10T12:00:00Z</updated>
<author><name>Alex</name></author>
<entry><title>RTX Spark laptops are on sale: $2,599 to start, $5,899 for 128GB</title><link href="https://measuredlocally.com/rtx-spark-laptop-prices"/><id>https://measuredlocally.com/rtx-spark-laptop-prices</id><updated>2026-10-10T12:00:00Z</updated><summary>Microsoft opened orders on Wednesday and five other makers follow on 16 October.  The 128GB version is the one that runs the big models, and I would wait for independent tests before buying one.</summary></entry>
<entry><title>Four DGX Sparks twice as fast at ordinary writing: the recipe, and a real session</title><link href="https://measuredlocally.com/dgx-spark-quad-recipe-speed"/><id>https://measuredlocally.com/dgx-spark-quad-recipe-speed</id><updated>2026-10-10T12:00:00Z</updated><summary>I moved my four DGX Sparks to a recipe that another owner published on GitHub, and ordinary writing roughly doubled on the same hardware and the same model.  The recipe is only a few weeks old, and these are one owner&#x27;s figures from one cluster.</summary></entry>
<entry><title>DGX Spark vs RTX Spark laptop: ports, price, and what joins up</title><link href="https://measuredlocally.com/dgx-spark-vs-rtx-spark-laptop"/><id>https://measuredlocally.com/dgx-spark-vs-rtx-spark-laptop</id><updated>2026-10-10T12:00:00Z</updated><summary>Microsoft&#x27;s RTX Spark laptop now costs less than NVIDIA&#x27;s own DGX Spark with the same memory.  The DGX Spark is the one that can be joined to other boxes, and I do not know yet how many of the recipes its owners share will work on Windows.</summary></entry>
<entry><title>NVIDIA&#x27;s 64GB DGX Spark at $4,999: what it runs</title><link href="https://measuredlocally.com/dgx-spark-64gb"/><id>https://measuredlocally.com/dgx-spark-64gb</id><updated>2026-10-03T12:00:00Z</updated><summary>NVIDIA announced a DGX Spark with 64GB of memory, half of what the original has, and it costs more than the full one did at launch.  The price is the part that stood out to me, and I could not find how NVIDIA measured the gain it claims from linking two or four together.</summary></entry>
<entry><title>M5 Ultra Mac Studio vs my four DGX Sparks: the same model</title><link href="https://measuredlocally.com/mac-studio-m5-ultra-vs-sparks"/><id>https://measuredlocally.com/mac-studio-m5-ultra-vs-sparks</id><updated>2026-10-03T12:00:00Z</updated><summary>MacStories tested GLM 5.3 Flash, the model my cluster runs, on a 256GB M5 Ultra Mac Studio, in a review I had missed.  One Mac Studio writes about as fast as my four DGX Sparks and my Sparks read a prompt 2.04 times as fast, but the comparison is a rough one.</summary></entry>
<entry><title>Anthropic&#x27;s report on GLM 5.3: what its tests found</title><link href="https://measuredlocally.com/glm-5-3-security-report"/><id>https://measuredlocally.com/glm-5-3-security-report</id><updated>2026-10-03T12:00:00Z</updated><summary>Anthropic published a security report on GLM 5.3, the family of open models my Sparks run.  By Anthropic&#x27;s count the model turned a known Chrome flaw into a working attack in 50 of 410 attempts, and as far as I could see on 3 October nothing had changed for anyone running it at home.</summary></entry>
<entry><title>Mac Studio M5 Ultra 512GB, and Apple&#x27;s 3x cluster claim checked</title><link href="https://measuredlocally.com/mac-studio-m5-ultra-512gb"/><id>https://measuredlocally.com/mac-studio-m5-ultra-512gb</id><updated>2026-09-28T12:00:00Z</updated><summary>Apple says its latest Mac Studio can hold 512GB of memory in one box, the same as my whole cluster of four DGX Sparks, and that a cluster of Studios runs AI up to 3x faster than one.  I looked up how Apple measured that, and I would wait for independent measurements before ordering.</summary></entry>
<entry><title>MiMo-V2.6 tops the open-weights board, and runs on DGX Sparks</title><link href="https://measuredlocally.com/mimo-v2-6-open-weights"/><id>https://measuredlocally.com/mimo-v2-6-open-weights</id><updated>2026-09-28T12:00:00Z</updated><summary>Xiaomi&#x27;s MiMo-V2.6-Pro sat at number one on Artificial Analysis&#x27;s open-weights leaderboard on 27 September, and DGX Spark owners had both versions running within days.  One owner reported 68.3 tokens a second on eight Sparks with speculative decoding, up from 17.8, and the early problem was tool use.</summary></entry>
<entry><title>Hugging Face Transformers now loads GGUF: the catch</title><link href="https://measuredlocally.com/gguf-in-transformers"/><id>https://measuredlocally.com/gguf-in-transformers</id><updated>2026-09-28T12:00:00Z</updated><summary>Hugging Face says its Transformers library can now load the GGUF files that llama.cpp made popular for running models at home.  For now it only keeps the file compressed for some Qwen models on Apple silicon, and I think it is most useful to people who work in Python.</summary></entry>
<entry><title>DeepSeek V4.1 Flash on 2 DGX Sparks, and the 2GB memory unlock</title><link href="https://measuredlocally.com/deepseek-v4-1-flash-two-sparks"/><id>https://measuredlocally.com/deepseek-v4-1-flash-two-sparks</id><updated>2026-09-19T12:00:00Z</updated><summary>One of the regular recipe authors on NVIDIA&#x27;s developer forum published a way to run DeepSeek V4.1 Flash on two DGX Sparks instead of four, with a second trick that applies to any headless Spark.  I had not run it myself, and the speeds and the one quality check were the author&#x27;s own.</summary></entry>
<entry><title>Ternary Bonsai 2: a 27B model in under 6GB</title><link href="https://measuredlocally.com/ternary-bonsai-2-27b"/><id>https://measuredlocally.com/ternary-bonsai-2-27b</id><updated>2026-09-19T12:00:00Z</updated><summary>Prism ML released a 27 billion parameter model built from Qwen 3.8 with every weight set to -1, 0 or +1, and says that brings it under 6GB.  The maker says it keeps 98.2% of the original&#x27;s scores, and on 19 September that was still the maker&#x27;s own number.</summary></entry>
<entry><title>Why a business bought 4 DGX Sparks: OpenFaaS&#x27;s month with them</title><link href="https://measuredlocally.com/openfaas-four-dgx-sparks"/><id>https://measuredlocally.com/openfaas-four-dgx-sparks</id><updated>2026-09-19T12:00:00Z</updated><summary>OpenFaaS, a UK software company, has written up why it spent about £10,000 on four DGX Sparks: customer data it will not send to a cloud provider, and red teaming its own products.  It also found two 4-bit builds of GLM 5.3 Flash running at different speeds, so two builds of the same model are not interchangeable.</summary></entry>
<entry><title>DeepSeek V4.1 Flash on 4 DGX Sparks: the 552B open model</title><link href="https://measuredlocally.com/deepseek-v4-1-flash-on-sparks"/><id>https://measuredlocally.com/deepseek-v4-1-flash-on-sparks</id><updated>2026-09-13T12:00:00Z</updated><summary>DeepSeek released a 552 billion parameter model with open weights on 10 September, and within a day an owner had it running on four DGX Sparks.  The speeds were similar to GLM 5.3 Flash on my own four, but they came from one owner&#x27;s first day on a four bit build that had not been checked for quality, as far as I could see.</summary></entry>
<entry><title>M5 Ultra Mac Studio wait times: 16 to 18 weeks for the full chip</title><link href="https://measuredlocally.com/mac-studio-m5-ultra-wait"/><id>https://measuredlocally.com/mac-studio-m5-ultra-wait</id><updated>2026-09-13T12:00:00Z</updated><summary>On 11 September Apple&#x27;s own store was quoting 16 to 18 weeks for any Mac Studio with the full M5 Ultra chip, which meant January.  The 512GB option, the one Mac that could hold a four bit build of DeepSeek V4.1 Flash, could not be ordered at all.</summary></entry>
<entry><title>NVIDIA&#x27;s NVFP4 builds of GLM 5.3 Flash: 18 checked for the fault</title><link href="https://measuredlocally.com/glm-5-3-flash-nvfp4-builds"/><id>https://measuredlocally.com/glm-5-3-flash-nvfp4-builds</id><updated>2026-09-13T12:00:00Z</updated><summary>NVIDIA published its own NVFP4 build of GLM 5.3 Flash on 9 September.  I checked 18 builds of that model and of Qwen 3.8 Flash Next for the fault I had found on my own cluster, and found it in three GLM builds, the most downloaded one among them, and in neither of NVIDIA&#x27;s own.</summary></entry>
<entry><title>RTX Spark: October launch, PAIR routing, a free RTX speed-up</title><link href="https://measuredlocally.com/rtx-spark-october-launch"/><id>https://measuredlocally.com/rtx-spark-october-launch</id><updated>2026-09-06T12:00:00Z</updated><summary>NVIDIA confirmed at IFA in Berlin that RTX Spark PCs would ship in October, and released PAIR, a free tool that spreads AI work across the PCs on a home network.  Nobody outside NVIDIA had benchmarked an RTX Spark, and I think the free speed-up for RTX cards mattered more on the day.</summary></entry>
<entry><title>NVIDIA buys Hugging Face for $12.93 billion: what it means</title><link href="https://measuredlocally.com/nvidia-buys-hugging-face"/><id>https://measuredlocally.com/nvidia-buys-hugging-face</id><updated>2026-09-06T12:00:00Z</updated><summary>NVIDIA agreed to buy Hugging Face for $12.93 billion.  For those of us running models at home that is our main supplier being bought by our hardware vendor, and I would make sure any model I rely on also lives on my own disk.</summary></entry>
<entry><title>Qwen3.8 Flash Next and GLM 5.3 Flash on DGX Spark: first speeds</title><link href="https://measuredlocally.com/qwen3-8-flash-next-on-sparks"/><id>https://measuredlocally.com/qwen3-8-flash-next-on-sparks</id><updated>2026-09-06T12:00:00Z</updated><summary>By 6 September owners on NVIDIA&#x27;s forum had Qwen3.8 Flash Next running on a single DGX Spark at 37 to 44 tokens a second, and one post had GLM 5.3 Flash on one box at 64.  The 64 came from a build squeezed to about 2 bits per weight, and I would want to see the quality numbers before copying it.</summary></entry>
<entry><title>Two DGX Sparks or a Mac Studio Ultra: the same 256GB, the numbers side by side</title><link href="https://measuredlocally.com/dgx-spark-vs-mac-studio"/><id>https://measuredlocally.com/dgx-spark-vs-mac-studio</id><updated>2026-09-01T12:00:00Z</updated><summary>A buyer with two new DGX Sparks asked on Reddit whether to send them back now that Apple had announced the M5 Ultra Mac Studio.  I put my own cluster&#x27;s numbers next to what Mac Studio owners measure, and I think the answer depends on the work: short questions suit the Mac, and agents and long documents suit the Sparks.</summary></entry>
</feed>
