New large language models ship every week — frontier models, open-weight releases, fine-tuned variants, and benchmark improvements. Keeping up requires more than a headline feed; builders need context on capability shifts, pricing changes, and API availability.
This cluster covers new model announcements, open-weight releases, benchmark claims, and model update signals from frontier labs and open-source communities. Each signal carries a link to the primary source for hands-on evaluation.
Use this page to maintain a model watchlist: identify capability jumps, track open-weight availability, and decide which models deserve a test run against your current workloads.
Where will the next breakout startup come from? Benchmark’s full partnership weighs in on the main stage at TechCrunch Disrupt 2026. Save up to $200 before September 25 at 11:59 p…
modelproduct
Sep 21, 2026·TechCrunchOriginal source → Google’s AI-native Googlebook ties Gemini to the cursor, dictation, widgets, and other parts of the desktop experience.
toolmodelcoding
Sep 21, 2026·TechCrunchOriginal source → macOS 27: Workaround to avoid downloading AI models and save storage — from Hacker News front page
modelinfra
Sep 21, 2026·Hacker NewsOriginal source → Release: llm-keys-ui 0.1 This plugin solves a very specific problem. I've started using Codex Remote to run coding agents on various machines while controlling them from my phone.…
toolmodelagentproduct
Sep 20, 2026·Simon WillisonOriginal source → Google said Gemini had "acted appropriately" by ending each hack immediately.
toolmodel
Sep 19, 2026·TechCrunchOriginal source → In May, Gemini broke containment and hacked three different companies, but Google didn't disclose the incident until the Wall Street Journal approached the company. The hacks happ…
toolmodelsecurity
Sep 19, 2026·The VergeOriginal source → Gemini Hacked Three Companies in First Known Breakout by Google’s AI Gemini finally caught up on Felony Bench ! The hacks, which the company confirmed on Friday, occurred in May a…
toolmodel
Sep 18, 2026·Simon WillisonOriginal source → Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programm…
modelagent
Sep 18, 2026·ArXivOriginal source → Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while exist…
modelagentinfra
Sep 18, 2026·ArXivOriginal source → Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design,…
toolmodelagent
Sep 18, 2026·ArXivOriginal source → Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoid…
toolmodelinfraproduct
Sep 18, 2026·ArXivOriginal source → LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inh…
toolmodelagent
Sep 18, 2026·ArXivOriginal source → Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cann…
toolmodelinfra
Sep 18, 2026·ArXivOriginal source → We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder a…
toolmodelagentcodinginfra
Sep 18, 2026·ArXivOriginal source → Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions. Existing memory systems typically rely…
toolmodelagent
Sep 18, 2026·ArXivOriginal source →