[LIVE]OPERATOR FEED · 50 ITEMS

Pulse

What changed for operators. Hardware. Runtimes. Models. Pricing. What to do about it. Not generic news.

RUNTIME6d ago· 1 entities linked

Cline Desktop v0.0.9: Universal macOS Build

The latest Cline Desktop release (v0.0.9) introduces a universal binary for macOS, eliminating the need to choose between Apple Silicon and Intel architectures. This update simplifies deployment by ensuring native performance on both hardware types without manual intervention. Th...

RUNTIME1w ago· 1 entities linked

llama.cpp b10273: Sampler Context Size Change

In version b10273 of llama.cpp, the default context size for history-based samplers has been standardized at 64. This change resolves an earlier issue where samplers were initialized before the complete llama_context was available, leading to uncertainty about the resolved contex...

RUNTIME1w ago· 1 entities linked

llama.cpp b10251 adds MTP support for GLM-4.7-Flash

The latest release of llama.cpp (b10251) introduces Multi-Tenant Plugin (MTP) support for the GLM-4.7-Flash model, enhancing compatibility and performance in multi-user environments. This update is crucial for operators running local AI services on consumer hardware such as Apple...

RUNTIME1w ago· 1 entities linked

koboldcpp v1.118.1 fixes multi-GPU issues and adds Qwen3TTS support

The latest release of koboldcpp (v1.118.1) addresses several critical bugs, including a fix for incorrect GPU selection in multi-GPU setups which could lead to performance degradation or service interruptions. Additionally, the update supports qwen3tts aliases for all voices and...

RUNTIME1w ago· 1 entities linked

koboldcpp v1.118: RPC compatibility and multi-GPU fixes

The latest release of koboldcpp (v1.118) introduces several critical updates for local AI operators running on consumer hardware. Notably, it aligns the RPC behavior with upstream standards, enabling seamless interoperability between llama.cpp RPC servers and clients alongside ko...

RUNTIME1w ago· 1 entities linked

llama.cpp b10215 adds Intel GPU driver check for Windows stability

The latest release of llama.cpp (b10215) introduces a critical update for users running on Windows with Intel GPUs. This version includes a driver version check to mitigate crashes that occur on certain Intel GPU drivers, specifically those prior to 32.0.101.8860. For operators r...

RUNTIME1w ago· 1 entities linked

llama.cpp b10199 adds inp embd support for next token generation

The latest release of llama.cpp (b10199) introduces significant server-side enhancements including the ability to generate the next token using input embeddings and sampled tokens. This update enables more sophisticated control over text generation processes, allowing operators t...

RUNTIME1w ago· 1 entities linked

llama.cpp b10189 Removes Custom CPU Op for M3 Graph

The latest release of llama.cpp (b10189) removes a custom CPU operation from the M3 graph, replacing it with standard operations. This change ensures better compatibility and performance across different hardware configurations without relying on proprietary optimizations. Operat...

MODEL2w ago· 1 entities linked

llama.cpp b10173 adds Laguna-S-2.1 LLM support

The latest release of llama.cpp (b10173) introduces support for the Laguna-S-2.1 large language model, enhancing capabilities for local AI deployments on consumer hardware like Apple Silicon and RTX GPUs. This update also includes binaries optimized for macOS ARM64 and Intel x64...

RUNTIME2w ago· 1 entities linked

ollama v0.32.5 fixes MLX Metal bug for NVFP4 models

ollama version 0.32.5 addresses a critical issue with MLX Metal that could degrade output quality, particularly affecting the Laguna model in the NVFP4 family. This update ensures local AI operators using these models on Apple Silicon hardware maintain high-quality outputs withou...

MODEL2w ago· 1 entities linked

llama.cpp b10142 Adds MiniMax-M3 Vision Support

The latest release of llama.cpp (b10142) introduces preliminary support for the vision tower in MiniMax-M3, a text-to-image model. This update includes components like mmproj and clip graph integration, enhancing the model's visual capabilities without sparse attention or MTP hea...

RUNTIME2w ago· 1 entities linked

ollama v0.32.4 adds Laguna support and Qwen3 fixes

ollama v0.32.4 introduces significant improvements for Apple GPU users by supporting the MLX engine on Laguna, enhancing performance and compatibility with Apple's hardware. Additionally, this release addresses a critical issue in Qwen3 MoE decoding, ensuring smoother operation a...

RUNTIME2w ago· 1 entities linked

sglang v0.5.16 introduces DSpark speculative decoding and Inkling support

The latest sglang release includes a significant update with the introduction of DSpark, a confidence-driven speculative decoding algorithm that enhances performance by dynamically sizing verify windows based on draft confidence rather than fixed lengths. This feature reaches imp...

RUNTIME2w ago· 1 entities linked

ollama v0.32.3: Improved GPU Support and Model Downloads

ollama v0.32.3 introduces several critical improvements for local AI operators running on consumer hardware. The release fixes model download stalls, enhances integrations with Claude Code Channels and Anthropic thinking streams, and ensures Hermes Desktop respects the `--force-b...

RUNTIME3w ago· 1 entities linked

Cline Desktop v0.0.2: Automatic Updates and macOS Support

The latest release of Cline Code for macOS introduces automatic update functionality that checks for new versions every two hours and downloads them in the background, prompting users to restart with a single click. This ensures seamless updates without manual intervention. The d...

RUNTIME3w ago· 1 entities linked

Cline CLI v3.0.45 cuts install size and adds Kimi K3

The latest Cline CLI release significantly reduces the installation footprint by making Claude Code and Codex providers optional dependencies, shrinking the global npm package from approximately 640MB to just over 285MB. This update also introduces support for the Kimi K3 model w...

HARDWARE3w ago· 3 entities linked

RTX 50 SUPER reportedly on hold: 3GB GDDR7 costs $60-70 vs $20 for 2GB

NVIDIA has reportedly told board partners to hold the GeForce RTX 50 SUPER launch. Per VideoCardz (2026-07-17), at least one AIB already has production-ready cards physically in hand, but NVIDIA has provided no new release schedule. The cited cause: 3GB GDDR7 packages currently r...

RUNTIME3w ago· 1 entities linked

Ollama 0.32.0: bare CLI now launches an agent; legacy models flagged deprecated

Ollama v0.32.0 (released 2026-07-11) changes default CLI behavior: running bare `ollama` now opens an interactive agent session — "Chat, Code, & Work" — that chats with models, writes code, searches the web, and delegates tasks. The release notes' example shows the agent defa...

RUNTIME3w ago· 1 entities linked

LM Studio ships Bionic: agent app for open models on Mac and Windows

LM Studio released Bionic on 2026-07-16, a separate Mac/Windows app that runs open models as agents for coding and knowledge work. Code projects attach to a local folder; the agent investigates, edits, and debugs with inline diffs and agentic code search. Work projects handle doc...

MODEL3w ago· 1 entities linked

Thinking Machines releases Inkling: 975B/41B-active MoE, Apache 2.0

Thinking Machines Lab released Inkling on July 15, 2026 — its first model, open-weight under Apache 2.0 with weights on Hugging Face at launch. It is a Mixture-of-Experts transformer: 975B total parameters, 41B active (6 of 256 routed experts plus 2 shared), context window up to...

MODEL3w ago· 2 entities linked

Kimi K3: 2.8T-param MoE announced, open weights due July 27

Moonshot AI announced Kimi K3 on 2026-07-16: a 2.8-trillion-parameter MoE with 896 experts, 16 active per token (~1.8% of the pool), a 1M-token context window, native vision, and always-on reasoning. The architecture pairs Kimi Delta Attention (hybrid linear attention) with "Atte...

RUNTIME3w ago· 1 entities linked

ollama v0.32.1: Improved multi-turn reasoning and bug fixes

ollama v0.32.1 introduces enhanced multi-turn reasoning capabilities with Gemma 4, offering more reliable tool-response continuations for complex queries. This update also resolves a persistent MLX model cache leak that could cause memory issues across requests and improves cache...

RUNTIME4w ago· 1 entities linked

vLLM v0.25.1 Fixes FFmpeg Dependency and RMSNorm Quant Fusions

The latest patch release of vLLM (v0.25.1) addresses two critical issues impacting local AI deployments: it defers runtime errors for missing system FFmpeg when importing TorchCodec, preventing startup blocks even if TorchCodec isn't used; and it guards mixed-dtype allreduce RMSN...

MODEL4w ago· 1 entities linked

llama.cpp b9993 Adds Support for Hy3 Model Architecture

The latest release of llama.cpp (b9993) introduces support for the Hy3 model architecture, which is based on Tencent's Hunyuan 3. This new addition includes a MoE decoder stack with per-head Q/K RMSNorm and an advanced router mechanism. Local AI operators should be aware that thi...

RUNTIME1mo ago· 1 entities linked

llama.cpp b9982: Per-Request Reasoning Budget Tokens Now Respected

In the latest release of llama.cpp (b9982), a critical bug has been fixed where per-request reasoning budget tokens were being ignored in chat completions. Previously, only server-level defaults and Anthropic-style aliases were considered, leading to potential misconfiguration or...

RUNTIME1mo ago· 1 entities linked

vLLM v0.25.0: Model Runner V2 as Default and PagedAttention Removal

The latest release of vLLM (v0.25.0) introduces significant changes that affect local AI operators running dense models. The new Model Runner V2 is now the default execution path, bringing support for EVS, real-time embeddings, prefix caching for Mamba hybrid models, multimodal-p...

RUNTIME1mo ago· 1 entities linked

LM Studio 0.4.19 makes Engine Protocol the default

LM Studio 0.4.19 (2026-07-07) turns the Engine Protocol on by default for all stable-channel users. llama.cpp options that previously lived in separate menus - chat template, CPU thread count, speculative decoding - now sit under Load Parameters. Behavior of loaded models does no...

RUNTIME1mo ago· 1 entities linked

vLLM 0.24.0 breaking change: CUDA_VISIBLE_DEVICES no longer set

vLLM v0.24.0 (2026-06-29) stops setting CUDA_VISIBLE_DEVICES internally. GPU selection moves to a new device_ids argument, with the old ROCm behavior on a deprecation window rather than removed outright. Existing multi-GPU launch configs that relied on vLLM managing device visibi...

RUNTIME1mo ago· 2 entities linked

Ollama 0.31.1: Gemma 4 up to ~90% faster on Apple Silicon

Ollama v0.31.1 (2026-06-30) enables multi-token prediction (MTP) in its MLX engine for Gemma 4 on Apple Silicon, reporting up to ~90% faster average generation in a coding-agent benchmark. MTP is on by default and auto-tuned - Ollama decides how many tokens to draft at runtime, o...

HARDWARE1mo ago· 2 entities linked

AMD Ryzen AI Halo 128GB mini-workstation hits retail at $3,999

AMD's first-party 128GB unified-memory Strix Halo box (Ryzen AI Max+ 395) reached retail on 2026-07-10 at $3,999 via Micro Center, in Windows and Linux SKUs. It is positioned directly against NVIDIA's DGX Spark at roughly $700 less, and AMD claims the 128GB unified pool can host...

MODEL1mo ago· 2 entities linked

Tencent Hunyuan 3.0 goes Apache 2.0 - EU/UK exclusion dropped

Tencent officially released Hunyuan 3.0 (295B MoE, 21B active) on 2026-07-06 under Apache 2.0, with an official FP8 quantization aimed at local deployment. The license reversal matters as much as the weights: the April preview excluded EU, UK, and South Korea users, and the Apach...

RUNTIME1mo ago

llama.cpp ships b9718

llama.cpp cut release b9718 on 2026-06-19. Release notes excerpt: "server : consolidate slot selection into get_available_slot (#24755) Absorb get_slot_by_id logic into get_available_slot so slot selection is handled by a single function call. When a specific slot id is requested...

RUNTIME1mo ago

llama.cpp ships b9721

llama.cpp cut release b9721 on 2026-06-19. Release notes excerpt: "sync : ggml **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9721/llama-b9721-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLE...

RUNTIME1mo ago

llama.cpp ships b9722

llama.cpp cut release b9722 on 2026-06-19. Release notes excerpt: "server: fix non-bound n_discard value (ctx shifting) (#24786) * server: fix non-bound n_discard value * Update tools/server/server-context.cpp Co-authored-by: Georgi Gerganov --------- Co-authored-by: Georgi Gerga...

RUNTIME1mo ago

llama.cpp ships b9714

llama.cpp cut release b9714 on 2026-06-19. Release notes excerpt: "server: add "X-Accel-Buffering": "no" header to streaming endpoints (#24774) * server: add "X-Accel-Buffering": "no" header to streaming endpoints This header tells Nginx (as a reverse proxy) to NOT buffer respons...

RUNTIME1mo ago

llama.cpp ships b9715

llama.cpp cut release b9715 on 2026-06-19. Release notes excerpt: "Ggml/cuda col2im 1d (#24417) * cuda: add GGML_OP_COL2IM_1D, follow-up to the CPU op * cuda: col2im_1d use fast_div_modulo for the index decomposition * cuda: col2im_1d tighten supports_op, type match and contiguou...

RUNTIME1mo ago

llama.cpp ships b9716

llama.cpp cut release b9716 on 2026-06-19. Release notes excerpt: "mtmd: add batching support for internvl (#24775) **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9716/llama-b9716-bin-macos-arm64.tar.gz) - macOS Apple Silic...

RUNTIME1mo ago

ComfyUI ships v0.25.1

ComfyUI cut release v0.25.1 on 2026-06-18. Release notes excerpt: "* [Partner Nodes] feat(Kling): add support for Kling V3-Turbo model (https://github.com/Comfy-Org/ComfyUI/pull/14528). **Full Changelog**: https://github.com/Comfy-Org/ComfyUI/compare/v0.25.0...v0.25.1"

RUNTIME1mo ago

llama.cpp ships b9704

llama.cpp cut release b9704 on 2026-06-18. Release notes excerpt: "server : return HTTP 400 on invalid grammar (#24144) (#24154) Throw on grammar parse failure so the server returns HTTP 400 instead of silently dropping the constraint. Add a regression test for the invalid-gramma...

RUNTIME1mo ago

llama.cpp ships b9702

llama.cpp cut release b9702 on 2026-06-18. Release notes excerpt: "server: fix router args not being forwarded to child instances (#24760) **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9702/llama-b9702-bin-macos-arm64.tar....

RUNTIME1mo ago

llama.cpp ships b9701

llama.cpp cut release b9701 on 2026-06-18. Release notes excerpt: "mtmd: refactor preprocessor, add mtmd_image_preproc_out (#24736) * add mtmd_image_preproc_out * add dev docs * remove unused clip API * rm unused clip_image_f32_batch::grid * change preprocess() call signature **m...

RUNTIME1mo ago

llama.cpp ships b9703

llama.cpp cut release b9703 on 2026-06-18. Release notes excerpt: "server: (router) rework -hf preset repo (#24739) * server: temporary remove HF remote preset * rework remove preset.ini support * rm unused get_remote_preset_whitelist() * print warning * add docs * rm stray file...

RUNTIME1mo ago

ollama ships v0.30.10

ollama cut release v0.30.10 on 2026-06-17. Release notes excerpt: "## What's Changed * models: add Cohere2MoE model by @jmorganca in https://github.com/ollama/ollama/pull/16670 * llama: update llama.cpp to b9672 by @pdevine in https://github.com/ollama/ollama/pull/16775 **Full Ch...

RUNTIME1mo ago

llama.cpp ships b9691

llama.cpp cut release b9691 on 2026-06-17. Release notes excerpt: "ggml-cpu: Conditionally enable power11 backend based on compiler support (#24687) * ggml: Conditionally enable power11 backend based on compiler support Guard POWER11 backend creation behind a compiler flag check...

RUNTIME1mo ago

llama.cpp ships b9692

llama.cpp cut release b9692 on 2026-06-17. Release notes excerpt: "mtmd: llava_uhd should no longer use batch dim (#24732) **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9692/llama-b9692-bin-macos-arm64.tar.gz) - macOS Appl...

RUNTIME1mo ago

llama.cpp ships b9688

llama.cpp cut release b9688 on 2026-06-17. Release notes excerpt: "server: (router) add model management API (#23976) * wip * server: (router) add SSE realtime updates API * nits * wip * add download API * add download api * update docs * add delete endpoint * fix std::terminate...

RUNTIME1mo ago

llama.cpp ships b9689

llama.cpp cut release b9689 on 2026-06-17. Release notes excerpt: "metal : add f16 and bf16 support for concat operator (#24724) * metal : add f16 and bf16 support for concat operator Extend the Metal backend concat operator to support f16 and bf16 tensor types in addition to the...

RUNTIME1mo ago

llama.cpp ships b9690

llama.cpp cut release b9690 on 2026-06-17. Release notes excerpt: "metal : implement rope_back operator (#24725) Reuse existing rope kernels with a function constant to toggle forward/backward rotation, avoiding duplicate kernel code. Assisted-by: pi:llama.cpp/Qwen3.6-27B **macOS...

RUNTIME1mo ago

llama.cpp ships b9678

llama.cpp cut release b9678 on 2026-06-17. Release notes excerpt: "opencl: optimize mul_mat_f16_f32_l4 for decode (#24504) **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9678/llama-b9678-bin-macos-arm64.tar.gz) - macOS Appl...

RUNTIME1mo ago

llama.cpp ships b9680

llama.cpp cut release b9680 on 2026-06-17. Release notes excerpt: "ci: fix vulkan docker images (#24595) * Update vulkan-shaders-gen.cpp * Update vulkan-shaders-gen.cpp add comment describing code change intention * Update vulkan-shaders-gen.cpp fix potential UB **macOS/iOS:** -...

[end of feed] · operator pulse · runlocalai.co/pulse