BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT: What are LLM benchmarks actually measuring?
// espacio-latente.com / radar
Píldora diaria de lo que se mueve en IA: laboratorios, papers, blogs de referencia y releases relevantes, resumido y con link al original. Se actualiza dos veces al día. · ← volver a espacio-latente.com · archivo · RSS
BenchMIRT: What are LLM benchmarks actually measuring?
langchain==1.4.0a4
langchain==1.4.0a3
v1.3.0
How AI-native companies turn workflows into operating capability
Path to Astra: critical capabilities and frontier safeguards
Healthcare organizations can now connect EHR and additional industry data to ChatGPT
Mapping global methane emissions from space with deep learning
Quoting Rick Brewster
Claude Fable 5.1 made me a really nice animated pelican
Codex bundles LibreOffice
GeoJSON Map Viewer
Quoting Tarn Adams
datasette-mcp 0.2
Python 3.15.0 candidate 2 is here!
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
Medical Causal Hypothesis Verification with Large Language Models
Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents
OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
Life Operators: a self-evolving framework for multiscale life modelling
Auditing Harness Tampering in Self-Improving Agents
Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops
KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training
Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
Do General NLP Embeddings Capture Ontological Reasoning?
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
Proactive cyber defense for governments and enterprises
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
OpenAI Astra and Looped Transformers
[AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens
Amazon’s AI assistant can now spot fake emails from the company
Researchers fear safety disaster ahead of OpenAI’s Astra release
The Trump administration is supporting OpenAI in the NYT copyright lawsuit
Google is sending MrBeast into the wilderness, armed with AI
OpenAI accused of ‘aiding and abetting’ Tumbler Ridge mass shooting in dozens of new lawsuits
NYC bans AI use for students until they reach high school
Google needs Hollywood more than the studios need AI
Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work
OpenAI delayed its new model’s development after the Hugging Face hack
Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’
We’re ‘dangerously close’ to dead internet theory, says Pangram’s CEO
US government sides with OpenAI on issue of training LLMs on copyrighted material
Wonderful more than doubles its valuation to $5B in under 6 months
India’s richest man now wants to turn aging computers into AI-ready PCs
HiddenLayer nabs $100M as enterprises rush to secure their AI deployments
PSA: Amazon’s shopping AI can now tell you if that message is a scam
Adobe acquires Indian market intelligence startup Rilo
OpenAI faces 30 more lawsuits tied to Tumbler Ridge shooting
AfterQuery reportedly becomes Y Combinator’s fastest-ever unicorn, now valued at $3.2B
OpenAI’s Astra model is on the way — and very good at breaking into computer systems
Google’s Android update tackles motion sickness, accessibility, and more
Gemini 3.8 Flash and 3.8 Flash Cyber
A Note from LWN
Paint.net 5.2 alpha now runs on Linux
I Don't Have a Smartphone
Mistral now trains on user input by default, except on enterprise tier
Biggest dark matter detector spots a single weird particle
Exit the Cave
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
SteamdDB Joins Nexus Mods
Firefox's AI Switch Is Off. Telemetry Isn't
Aging Brains Blend Memories Together Instead of Just Forgetting Them
Commodore 64 released September 1, 1982
v3.7.0
Construir este día costó 69.095 tokens de entrada y 0 de salida en 73 llamadas a modelos — unos 0,0691 $.