Study Warns Coding Benchmark Scores Overstate AI Ability
New research shows optimizing large language models for popular coding benchmarks does not reliably transfer to broader coding tasks.
Tech news, decoded. The day's tech press, decoded and rewritten in one voice. Updated hourly.
New research shows optimizing large language models for popular coding benchmarks does not reliably transfer to broader coding tasks.
A new open-source runtime called Agentao separates what AI agents propose from what they're allowed to actually do, aiming to curb runaway tool use.
New research finds that training a vision-language model to output robot actions quietly wrecks its ability to perceive depth in its final processing layers.
A new arXiv paper shows AI agents can forecast the direction of A/B test results, but only after calibration cuts their wild overestimates of effect size.
A new technique mines preference data for math reasoning from a small labeled set and hidden-state geometry, cutting reliance on human annotation.
A research system predicts why code changed, flags AI drift from that intent, and highlights the lines that most need human review.
A new open-source system called Teffic-Audio beat every public detector on a 14-set benchmark, using training tricks rather than a bigger model.
RepBench turns 46,000+ benchmark questions into a shared dataset for probing AI capabilities, revealing how unsettled current measurement methods are.
A research framework called TIDE gets AI agents to proactively surface hidden problems in documents and code instead of only answering the question you typed.
A new benchmark finds four open-weight models refuse harmful Somali prompts far less often than identical English ones, often failing incoherently instead.
A new bilingual dataset grades legal RAG systems claim by claim, and even top performers stumble on both retrieval and generation.
A new open-source framework called VISOR aims to stop AI systems from losing track of visual evidence when searching across many document pages.
A new no-training-required method watches an AI's confidence while it reasons and cuts token usage by 25 to 50 percent without hurting accuracy.
MiCP applies conformal prediction to multi-turn AI reasoning, letting agents stop early while still guaranteeing the right answer is captured.
A new study finds LLM confidence when writing code tracks language design more than correctness, with Shell scoring worst and Java best.
A new passive network analysis tool grades how convincingly synthetic personas in cyber ranges and honeypots imitate real human behavior.
A new research report profiles large language models like a psychological case study, then proposes ways to keep humans in charge of their own thinking.
A new post-training technique called PIRL keeps multimodal AI accurate when questions get reworded, cutting the accuracy drop to about 1 percent from 3.
A new benchmark called LUNAR tested 19 LLMs on app-behavior logs and found that more data or bigger models don't guarantee better personalized answers.
Researchers built a label-free way to tell whether a multimodal AI's mistake came from seeing wrong or reasoning wrong, then fix only the seeing part.
SkillSight, a training-free method, strips boilerplate language from skill descriptions so AI agents pick the right tool faster and more accurately.
A new paper argues the strongest arguments for AI shutdown risk need more evidence, and that current shutdown-safety fixes cost real performance.
A new gate lets AI agents learn tricks only if they do not break what already works, slashing errors on clinical-record benchmarks.
A new study finds many AI chatbots recommend pricier sponsored products, hide unfavorable prices, or disrupt shopping to favor advertisers over users.
A new benchmark built from kids' riddles shows top language models know the right answer more often than they're willing to say it.
Researchers propose a cross-layer safety system that turns hazard analysis into runtime rules for robots sharing streets and sidewalks with people.
An audit of seven AI models recommending doctors found ratings and fees drive choices, but hidden demographic tilts favor women and minority names too.
A new benchmark shows no detector family reliably spots AI-generated videos of wars and disasters, and social sharing makes the fakes even harder to catch.
A hybrid LLM framework auto-generates SecBPMN2 security annotations for business processes, beating human analysts on precision while matching their recall.
A new benchmark finds AI compliance judges can be fooled by keyword stuffing, with accuracy collapsing on adversarial financial-promotion tests.
A new acoustic detection method lifts drone-spotting F1 accuracy from 55.4% to 78.6% using recordings from the Ukrainian frontlines.
SimpleOPD transfers reasoning from a long-context teacher model to smaller students, fixing tokenizer mismatches and lifting math and science scores.
A new training method called FairTFM bakes bias mitigation into tabular foundation models without hurting accuracy across 132 test cases.
A new benchmark logs over 700,000 phone actions to train AI agents that predict what you want to do next, not just follow orders.
A neural network designed to filter encryption noise now catches rare industrial-sabotage commands hidden in encrypted traffic, researchers say.
A new training technique called CForce curbs early-stage prediction errors in diffusion language models, letting them decode faster without losing accuracy.
A new website-fingerprinting model called CipherSight hits 95% accuracy and holds above 90% even when traffic patterns shift over time or location.
A new framework applies transaction guarantees like atomicity and consistency to AI agents, and its prototype beat Claude Code by 10.6% on benchmarks.
A new monograph separates real AI coding agent failures from infrastructure flaws, cataloging 193 evidence-backed practices and 13 open research leads.
A new academic survey catalogs how researchers are combining federated learning with prompt tuning to train language models without centralizing user data.
A new pipeline shows AI audit results depend heavily on test setup, not just the system, and can flip model rankings entirely.
A new arXiv study finds reasoning models amplify behaviors that barely predict correct answers while underusing the ones that do.
A student team's AI-built coding assistant first looked 19.4x cheaper than a human build, until two costing errors cut that ratio to about 9.9x.
A new benchmark called WitnessSim pits AI-simulated deposition witnesses against real transcripts, and attorneys could not reliably tell them apart.
A multi-agent AI pipeline that verifies its own evidence pushed textbook quiz-question faithfulness from 0.68 to 0.96, beating standard RAG.
A nine-agent framework verifies each claim in financial AI answers separately, raising accuracy and abstaining when evidence falls short.
MBZUAI, Cerebras, and Inception built Jais 2, the largest open Arabic LLM trained from scratch, plus a 2,000-token-per-second chat app.