Mistral Large 4 opens as an API preview, with weights promised by October’s end
Mistral launched an image-capable Large 4 API preview on October 6, while its promised public weights and license remain pending.
Muse’s memory report raises profiling questions, while Meta describes separate user environments
Reporting on Muse’s relationship files raises privacy questions, while Meta’s public disclosure describes per-user environments and sanitized product learning.
OpenAI’s Ironclad research improves workflow scores, without announcing a customer agent
OpenAI reports GPT-6 Astra met 55% of criteria across 11 Ironclad research tasks, up from 41.6% for GPT-5.6 Sol.
Anthropic expands cyber access through three tiers, including Mythos 5.1
Anthropic’s expanded Cyber Verification Program offers qualifying defenders tiered access to advanced Claude cyber capabilities, including Mythos 5.1.
Utah authorizes initial AI acne prescriptions, starting with two-physician approval
Utah’s one-year Nolla pilot permits initial topical acne prescriptions under staged oversight, with two physicians approving every launch-stage order.
Google releases EmbeddingGemma 2 for local search across words, pictures, and sound
Google’s EmbeddingGemma 2 maps text and media into a shared search space, with a quantized full configuration measured at about 567 MB of active phone RAM.
OpenAI launches a Decisions API that returns choices instead of prose
OpenAI’s Decisions API beta returns bounded decisions from text and images at $0.10 per million input tokens, with no output-token charge.
SemiAnalysis measures a fivefold subscription-value gap, with a workload-specific denominator
SemiAnalysis estimates roughly fivefold API-equivalent subscription value for one Claude-versus-OpenAI comparison, without measuring completed coding work.
TasteVal reports cheaper experimental search, without demonstrating scientific invention
TasteVal reports Opus 5.5 reached an expert benchmark reference with a 2.30 compute multiplier across eight constrained experimental-research tasks.
ProjectDiscovery demonstrates a poisoned model turning tool access into credential theft
ProjectDiscovery reports a trigger-linked model backdoor that retrieved a credential-collecting payload through Codex CLI in a controlled demonstration.
Queen joins chess expertise to language, reaching an estimated 2697 rating
Princeton’s Queen system reports an estimated 2697 chess rating after iterative explanation training, without proving its prose faithfully traces move selection.