Issue 001

AI Engineering Digest

Sourced briefing for engineers

Subscribe via RSS: feed.xml — paste that URL into a reader. No account, no email list.

Anthropic Inference Hooks (inline DLP, beta)

Claude Enterprise can POST the conversation to a customer/vendor HTTPS security server and wait for allow/deny before inference. One hook covers claude.ai, Cowork, and Claude Code; not Bedrock, Google Cloud, Platform API orgs, or voice. Shadow mode, rollout %, fail-open/fail-closed. Blog says “signed WebSocket”; docs say HTTPS POST — trust the docs for protocol.

Google Agent Platform evals GA

Same metric engine for offline experiments and online monitors on live Cloud Trace. 20+ prebuilt metrics (task success, tool use, trajectory, safety, grounding), adaptive rubrics, custom code/LLM metrics. You pay model-judge + storage; data stays in the project.

LangSmith Tuned Evaluators (Perceived Error)

Managed, versioned judges on production threads. First evaluator is Perceived Error (user correction, unresolved outcome). Docs: beta; US Plus and Cloud Enterprise; billed only when feedback attaches. Blog “up to 82%” cheaper / “98%” partner figures are vendor numbers. vendor

Anthropic GA: computer use, Skills API, Files API

Computer use adds browser use + multi-action turns and is HIPAA-eligible under Anthropic’s BAA. Files API: auto-expiration, “5x higher rate limits,” vendor 1 TB/org. Skills/Files also on Microsoft Foundry; Vertex “coming soon.” Customer 32→13 min / ~30% cost quote is not independently verified. vendor

OWASP GenAI LLM Top 10 2026

Updated rankings, incident-grounded research, mappings to NIST, MITRE ATLAS, CWE, Agentic Top 10. Project page says Aug 3; Foundation/GitHub banner says Aug 4. Companion arXiv (not the official list) says expert-vs-incident agreement is weak (κ ≈ 0.20).

Snowflake Cortex dynamic model routing (preview soon)

Cortex AI Gateway picks the cheapest approved model that meets a quality bar per agent step; logged, allowlisted, residency-aware. Status: “PrPr soon,” not GA. DeepSeek-V4-Flash 0731 private preview; GLM-5.3 “coming soon.” Internal 3× / 25% token claims are vendor-only. vendor TechTarget said gateway launched July 2025; Snowflake’s 2026 text says announced July 2026.

Ollama 0.32.7–0.32.15

0.32.7 (Aug 10): ollama run muse-glimmer:30b-mlx (MLX-first on Apple Silicon). 0.32.9 (Aug 11): nemotron-3.5-lightning (30B MoE / 3B active). 0.32.12 (Aug 14): qwen3.8:27b and qwen3.8:27b-mlx. 0.32.15 (Aug 19): claimed TTFT ~995→~524 ms (Ollama’s bench). vendor

MLX 0.32.1

Metal gemv_wide; NAX attention unroll; head dim 96/72; higher qmv batch on M5; GGUF metadata/offset fixes; Metal 4.1 / macOS 27 compiler request; zero-copy CPU import on unified memory.

No new mlx-lm tag after v0.31.3 (2026-04-22). mlx-swift latest tag 0.31.6 (2026-07-02) is just before the window. WWDC26 session 232 is the official four-layer recipe (June, not August news): developer.apple.com/videos/play/wwdc2026/232.

xAI Grok 4.6

Long-running agents + visual work. Same-day Cursor, Grok Build, SpaceXAI API, OpenRouter/Vercel/Cloudflare. $2 / $6 per 1M in/out (fast variant 2×). Docs: grok-4.6, 500k context, cutoff 2026-02-01. Vendor evals (High): AA Intelligence Index 61, CursorBench 3.2 69.9%, DeepSWE 1.1 65.9%. vendor Copilot Aug 14; Bedrock GA Aug 19 at $2 / $0.50 cached / $6. “Half the price of rivals” is not on the launch post.

Google Gemini 3.7 Flash

Workhorse for coding/agents, three weeks after 3.6 Flash. Intro $0.75 / $3.75 per 1M through 2026-12-31; then $1.50 / $7.50. Official vs 3.6 Flash: FrontierCode 1.1 Main 43.6 vs 34.4; DeepSWE v1.1 65.3 vs 49.0; WebDev Arena Elo 1588 vs 1538. vendor

DeepSeek-V4-Pro GA + Flash vision exp

deepseek-v4-pro now serves V4-Pro-0813; thinking effort low/high/max; Responses API. Peak/off-peak from 16:00 UTC Aug 16. Official table (per 1M): Pro miss-in $0.66 off-peak / $1.32 peak, out $1.98 / $3.96; Flash $0.22/$0.44 in, $0.66/$1.32 out. Peak 01:00–04:00 and 06:00–10:00 UTC. Aug 21: deepseek-v4-flash-vision-exp (1M context, 384K max out). Press “46× cheaper than Claude” is not on DeepSeek pages.

OpenAI: Ultrafast Sol + Sol promo (not new weights)

Cerebras-backed Ultrafast: GPT-5.6 Sol up to 14× Standard, vendor up to 750 out tok/s, limited preview. ChatGPT Aug 6: Plus/Pro updated Sol + effort slider; Free/Go default Luna. Aug 21: Sol API/credit “over 20%” down ~3 months; vendor short-context table now $4 / $20 through at least 2026-11-21. Family GA was 2026-07-09.

No new Microsoft-trained frontier LLM. Mistral Shieldstral (Aug 4) is a 3B safety classifier. Discarded: Fable 5.1 leaks; OpenAI Astra (named, not shipped). Qwen3.8-Max blog fetch returned a shell; treat the 2.4T open card in section 7 as the verified artifact.

1. Muse Glimmer coding agent via Ollama MLX

Install Ollama ≥0.32.7 (prefer 0.32.15). ollama run muse-glimmer:30b-mlx, then ollama launch pi --model muse-glimmer:30b-mlx (or claude/hermes). Treat 32 GB unified memory as a realistic floor. Why now: first Ollama path is MLX-only on Mac (Aug 10); 0.32.15 (Aug 19) updates MLX/llama.cpp and cuts their reported TTFT.

Not a laptop weekend project: DeepSeek V4 Flash local (LM Studio: ≥156 GB). Not “why now”: a new mlx-lm or mlx-swift tag (none in-window).

vLLM 0.27.0

Day-0 Kimi K3; Qwen3.5 dense/MoE; PyTorch 2.13 / Triton 3.7.1 (breaking). FlashAttention 4 on SM100; DeepSeek-V4 kernel/TTFT; Rust frontend gRPC.

SGLang 0.5.17

Day-0 Kimi K3 (2.8T LatentMoE, 1M context, native MXFP4) and MiniMax-H3 video+stereo-audio. Initial Rust frontend; SM90 FP8 MegaMoE for DeepSeek-V4. Breaking: helion 0.2.6→1.4; sglang.jit_kernel retired.

Transformers 5.15.0

First-class Muse Glimmer, GraniteMoeSWA, SKT A.X-K1/K2, Cosmos3 Edge. Native Mistral tekken in AutoTokenizer. Patch 5.15.1 on Aug 19.

LlamaIndex v0.14.24

Claude Sonnet 5 / Opus 5, GPT-5.6, Gemini 3.7 Flash default, MCP 2.x tools. Previous core tag was June 24.

Open WebUI v0.11.0

UI rebuild, admin-gated sub-agents, LDAP group sync, Anthropic passthrough, LiteLLM connection type. Security advisory: XSS in terminal HTML/math, chat-write auth, OAuth token-exchange hardening.

Kimi K3 (~2.8T, custom license)

LMSYS/SGLang: first open ~3T-class multimodal; hybrid 69 KDA + 24 MLA; LatentMoE 896 experts top-16; MoonViT3d; 1M context; native MXFP4. Day-0 SGLang then vLLM 0.27.0. License is Kimi K3, not MIT/Apache — read before MaaS.

DeepSeek-V4-Flash-0731 (MIT)

Official Flash, supersedes preview. DSpark in checkpoint. Vendor table: Terminal Bench 2.1 82.7, DeepSWE 54.4 vs preview 61.8 / 7.3. vendor vLLM recipe dates official Flash 2026-07-31.

DeepSeek-V4-Pro-0813 (MIT)

Official Pro. Vendor: HLE 42.7 / 60.0 (no tools / tools), Terminal Bench 2.1 87.9, DeepSWE 62.7. vendor reasoning_effort low|high|max. Chat template is Python in encoding/, not Jinja. HF ~1.7T. Card has no printed YYYY-MM-DD.

Qwen3.8-2.4T-A95B

2.4T total / 95B active; hybrid Gated DeltaNet + Gated Attention + MoE; native 262,144 context, extensible to 1,010,000. Text-only; thinking required. License not stated on the fetched card — do not assume Apache 2.0. Exact day medium confidence (August 2026). Official qwen.ai/blog/qwen3.8-max fetch returned a site shell.