Claude Enterprise can POST the conversation to a customer/vendor HTTPS security server and wait for allow/deny before inference. One hook covers claude.ai, Cowork, and Claude Code; not Bedrock, Google Cloud, Platform API orgs, or voice. Shadow mode, rollout %, fail-open/fail-closed. Blog says “signed WebSocket”; docs say HTTPS POST — trust the docs for protocol.
Same metric engine for offline experiments and online monitors on live Cloud Trace. 20+ prebuilt metrics (task success, tool use, trajectory, safety, grounding), adaptive rubrics, custom code/LLM metrics. You pay model-judge + storage; data stays in the project.
Managed, versioned judges on production threads. First evaluator is Perceived Error (user correction, unresolved outcome). Docs: beta; US Plus and Cloud Enterprise; billed only when feedback attaches. Blog “up to 82%” cheaper / “98%” partner figures are vendor numbers. vendor
Computer use adds browser use + multi-action turns and is HIPAA-eligible under Anthropic’s BAA. Files API: auto-expiration, “5x higher rate limits,” vendor 1 TB/org. Skills/Files also on Microsoft Foundry; Vertex “coming soon.” Customer 32→13 min / ~30% cost quote is not independently verified. vendor
Cortex AI Gateway picks the cheapest approved model that meets a quality bar per agent step; logged, allowlisted, residency-aware. Status: “PrPr soon,” not GA. DeepSeek-V4-Flash 0731 private preview; GLM-5.3 “coming soon.” Internal 3× / 25% token claims are vendor-only. vendor TechTarget said gateway launched July 2025; Snowflake’s 2026 text says announced July 2026.
Launch-day Muse Glimmer (tools + image). DeepSeek V4 Flash local: 284B MoE; they say plan for ≥156 GB RAM. Bionic itself shipped 2026-07-16 (older). Index dates: DeepSeek Aug 4, Muse Aug 17 (bodies undated).
Metal gemv_wide; NAX attention unroll; head dim 96/72; higher qmv batch on M5; GGUF metadata/offset fixes; Metal 4.1 / macOS 27 compiler request; zero-copy CPU import on unified memory.
Converted with mlx-vlm 0.6.8. python -m mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit. Optional MTP drafter. Hub “updated ~8 days ago” at research time, not an ISO first-publish date.
No new mlx-lm tag after v0.31.3 (2026-04-22). mlx-swift latest tag 0.31.6 (2026-07-02) is just before the window. WWDC26 session 232 is the official four-layer recipe (June, not August news): developer.apple.com/videos/play/wwdc2026/232.
Long-running agents + visual work. Same-day Cursor, Grok Build, SpaceXAI API, OpenRouter/Vercel/Cloudflare. $2 / $6 per 1M in/out (fast variant 2×). Docs: grok-4.6, 500k context, cutoff 2026-02-01. Vendor evals (High): AA Intelligence Index 61, CursorBench 3.2 69.9%, DeepSWE 1.1 65.9%. vendor Copilot Aug 14; Bedrock GA Aug 19 at $2 / $0.50 cached / $6. “Half the price of rivals” is not on the launch post.
Workhorse for coding/agents, three weeks after 3.6 Flash. Intro $0.75 / $3.75 per 1M through 2026-12-31; then $1.50 / $7.50. Official vs 3.6 Flash: FrontierCode 1.1 Main 43.6 vs 34.4; DeepSWE v1.1 65.3 vs 49.0; WebDev Arena Elo 1588 vs 1538. vendor
deepseek-v4-pro now serves V4-Pro-0813; thinking effort low/high/max; Responses API. Peak/off-peak from 16:00 UTC Aug 16. Official table (per 1M): Pro miss-in $0.66 off-peak / $1.32 peak, out $1.98 / $3.96; Flash $0.22/$0.44 in, $0.66/$1.32 out. Peak 01:00–04:00 and 06:00–10:00 UTC. Aug 21: deepseek-v4-flash-vision-exp (1M context, 384K max out). Press “46× cheaper than Claude” is not on DeepSeek pages.
Cerebras-backed Ultrafast: GPT-5.6 Sol up to 14× Standard, vendor up to 750 out tok/s, limited preview. ChatGPT Aug 6: Plus/Pro updated Sol + effort slider; Free/Go default Luna. Aug 21: Sol API/credit “over 20%” down ~3 months; vendor short-context table now $4 / $20 through at least 2026-11-21. Family GA was 2026-07-09.
No new Microsoft-trained frontier LLM. Mistral Shieldstral (Aug 4) is a 3B safety classifier. Discarded: Fable 5.1 leaks; OpenAI Astra (named, not shipped). Qwen3.8-Max blog fetch returned a shell; treat the 2.4T open card in section 7 as the verified artifact.
Install Ollama ≥0.32.7 (prefer 0.32.15). ollama run muse-glimmer:30b-mlx, then ollama launch pi --model muse-glimmer:30b-mlx (or claude/hermes). Treat 32 GB unified memory as a realistic floor. Why now: first Ollama path is MLX-only on Mac (Aug 10); 0.32.15 (Aug 19) updates MLX/llama.cpp and cuts their reported TTFT.
Use the first advertised stable tag (Aug 21). llama serve -hf meta-models/Muse-Glimmer-30B-GGUF; try DFlash draft flags from the HF blog. 32 GB+ preferred.
Standard model exception types; tighter tool-schema serialization. Also verified, not expanded: LiteLLM v1.97.0 (2026-08-16), CrewAI 1.15.17 (2026-08-20).
LMSYS/SGLang: first open ~3T-class multimodal; hybrid 69 KDA + 24 MLA; LatentMoE 896 experts top-16; MoonViT3d; 1M context; native MXFP4. Day-0 SGLang then vLLM 0.27.0. License is Kimi K3, not MIT/Apache — read before MaaS.
Official Pro. Vendor: HLE 42.7 / 60.0 (no tools / tools), Terminal Bench 2.1 87.9, DeepSWE 62.7. vendorreasoning_effort low|high|max. Chat template is Python in encoding/, not Jinja. HF ~1.7T. Card has no printed YYYY-MM-DD.
2.4T total / 95B active; hybrid Gated DeltaNet + Gated Attention + MoE; native 262,144 context, extensible to 1,010,000. Text-only; thinking required. License not stated on the fetched card — do not assume Apache 2.0. Exact day medium confidence (August 2026). Official qwen.ai/blog/qwen3.8-max fetch returned a site shell.
From 2 Aug 2026, AI Office and national authorities begin enforcing. Interactive systems must disclose AI; deepfakes labelled; AI-generated/altered content needs machine-readable marks. Systems on the market before 2 Aug 2026 have until 2 Dec 2026 for Article 50(2) marking. High-risk Annex III later (FAQ: 2 Dec 2027).
arXiv 2608.11693v2: PTX does not expose 5th-gen integer tensor-core .kind::i8 on sm_103a; CUTLASS skips INT8 UMMA for 103a; vLLM hard-errors; SGLang AOT INT8 GEMM stops at Sm90. Single paper, not vendor confirmation.