AI Brief

Research & papers

94 curated stories.

Research & papers · arXivPaper argues deployment rules shape multi-agent AI safetyResearch & papers · arXivSkillCenter paper proposes a source-grounded skill library for AI agentsResearch & papers · arXivPaper maps recursive self-improvement into autonomous research loopsResearch & papers · arXivPaper finds deterministic gates can catch tool-agent policy failuresResearch & papers · arXivPaper finds early signals for doomed LLM agent runsResearch & papers · arXivDataGovBench tests LLM data analysis on government open dataResearch & papers · arXivPaper tests evidence-linked multi-agent AI on biopsy reportsResearch & papers · arXivX-FEMR paper proposes token-level explanations for EHR foundation modelsResearch & papers · MarkTechPostNVIDIA HORIZON uses git worktrees to iterate hardware-design agentsResearch & papers · MarkTechPostNVIDIA ASPIRE turns robot debugging into a reusable skill libraryResearch & papers · arXivResearchers show coding-agent attacks can be spread across pull requestsResearch & papers · arXivResearchers propose real-time online safety monitoring for LLM outputsResearch & papers · arXivReContext improves long-context reasoning by replaying evidenceResearch & papers · arXivStudy finds AI agents say different things off the recordResearch & papers · arXivPaper says constraints can make coding-agent review scale betterResearch & papers · arXivResearchers propose hardware-enforced coordination for autonomous AIResearch & papers · arXivPACE proxy benchmark predicts agentic evaluation scores at under 1% of full costResearch & papers · MarkTechPostAlibaba Page Agent controls web interfaces through the DOMResearch & papers · arXivVera benchmark finds high attack success against production agent frameworksResearch & papers · arXivMicrosoft study tracks command-line AI coding-agent adoptionResearch & papers · Google ResearchGoogle AMIE research moves from diagnosis to disease managementResearch & papers · arXivAxDafny paper tests agentic code generation against formal verificationResearch & papers · arXivMARS paper uses text refusal directions to improve multimodal model safetyResearch & papers · arXivProtoPilot paper shows self-evolving agents for wet-lab protocolsResearch & papers · arXivWorldEvolver paper targets self-evolving world models for agentsResearch & papers · arXivResearchers identify entity-binding failures in tool-using agentsResearch & papers · arXivNew paper asks whose side an AI agent is onResearch & papers · arXivData-centre paper frames AI infrastructure as a many-body problemResearch & papers · arXivResearchers propose an immune-system model for agent safetyResearch & papers · arXivTandem reinforcement learning links language and action for agentsResearch & papers · arXivResearchers show prompt injection can manipulate AI resume screeningResearch & papers · arXivNew benchmark asks when combining frontier models actually helpsResearch & papers · arXivResearchers argue benchmarks miss much of collective model capabilityResearch & papers · The RegisterStudy warns medical diagnosis AIs can reveal training-data membershipResearch & papers · arXivNew paper maps how AI changes enterprise software user rolesResearch & papers · The VergeThe Atlantic makes AI music-training datasets searchableResearch & papers · arXivLedgerAgent proposes structured state for policy-adherent agentsResearch & papers · arXivMulti-LCB extends LiveCodeBench across programming languagesResearch & papers · arXivPaper probes what safety-aligned LLMs learn from mixed complianceResearch & papers · arXivTxBench-PP tests AI agents on preclinical pharmacologyResearch & papers · arXivSciRisk-Bench targets AI-for-science safety evaluationResearch & papers · Financial TimesFT says AI medical tools matched or surpassed doctorsResearch & papers · arXivRed-team paper finds sustained attacks still break frontier Anthropic modelsResearch & papers · OpenAIOpenAI details deployment simulation for model-risk forecastingResearch & papers · Financial TimesStudy finds Mistral more vulnerable to Russian disinformationResearch & papers · arXivPaper argues U.S. controls helped accelerate China open AI ecosystemsResearch & papers · The GuardianUK Nerve Lab uses AI to map how screen time affects childrenResearch & papers · NvidiaNvidia publishes the first AgentPerf infrastructure resultsResearch & papers · AnthropicAnthropic survey finds AI adoption outpacing public trustResearch & papers · arXivClinHallu diagnoses stage-by-stage hallucinations in medical AIResearch & papers · arXivAgentCyberRange tests frontier AI systems in realistic cyber rangesResearch & papers · arXivAFFORDANCE20Q probes AI reasoning about physical propertiesResearch & papers · arXivUniversal Manipulation Exoskeleton releases physical-AI data stackResearch & papers · Google DeepMindDeepMind and partners commit up to $10M to multi-agent safetyResearch & papers · The RegisterAI memory and personalization can increase sycophancyResearch & papers · arXiv / SciAgentArena researchersSciAgentArena tests AI agents on real scientific research workflowsResearch & papers · TechCrunchAI memory tools can degrade model performanceResearch & papers · Associated Press + MITMIT turns hidden hand motion into robot-training dataResearch & papers · The Register + CheckmarxAI-heavy teams ship vulnerable code at 3.4 times the rateResearch & papers · TechRadar + Linux FoundationEuropean employers expect AI to increase tech hiringResearch & papers · arXivActProbe spots robot-policy failures before they become visibleResearch & papers · arXivA self-evolving scientific agent discovers an interpretable fluid controllerResearch & papers · Unite.AI + SCAMMER4U researchersWeb agents leak sensitive data even after recognizing scamsResearch & papers · arXivPACE tries to stop self-improving agents from p-hacking themselvesResearch & papers · The Register + ForresterForrester finds enterprise agents still trapped in pilot modeResearch & papers · ServiceNow AI / Hugging FaceServiceNow expands EVA-Bench for real enterprise voice-agent workflowsResearch & papers · AgentScout / arXiv trackerAgent research week centers on self-evolving agents and governanceResearch & papers · arXivMIRAI tries to predict which research will matter years laterResearch & papers · AnthropicAnthropic maps AI-enabled cyber abuse across 832 banned accountsResearch & papers · Nature CommunicationsSuperARC benchmark says frontier models remain far from its AGI targetResearch & papers · Nature / Pediatric ResearchNature paper tests multimodal AI for a dangerous neonatal diseaseResearch & papers · University of TorontoResearchers demonstrate an AI worm that adapts as it spreadsResearch & papers · ChatPaper / arXivClinEnv benchmarks LLM agents as attending physicians over full inpatient staysResearch & papers · ChatPaper / arXivMOC paper targets message quality inside multi-agent AI systemsResearch & papers · EurekAlert / IRB BarcelonaIRB Barcelona uses generative AI to design cell-selective moleculesResearch & papers · OpenClaw / Hugging FaceOpenClaw releases a security dataset for agent skillsResearch & papers · Nature Machine IntelligenceNature Machine Intelligence paper links climate modes with machine learningResearch & papers · VentureBeat / arXivMeMo proposes memory models as an alternative to noisy enterprise RAGResearch & papers · arXivMAVEN shows tool-calling agents still struggle to generalizeResearch & papers · arXivEHRBench scales clinical-decision testing for medical LLMsResearch & papers · VentureBeat / arXivAutoTTS uses an AI-designed controller to cut reasoning-token use 69.5%Research & papers · EurekAlert / Annals of Family MedicineAI-assisted ultrasound helps Shanghai GPs spot carotid plaqueResearch & papers · VentureBeat + DatacurveDeepSWE challenges coding-agent leaderboards and benchmark trustResearch & papers · arXivAgentHijack tests computer-use agents against ordinary UI disruptionResearch & papers · Howard University + AWS coverageHoward University launches an AWS-powered AI networkResearch & papers · OpenAIOpenAI model disproves a decades-old discrete geometry conjectureResearch & papers · Google DeepMindGoogle DeepMind highlights Antigravity 2.0 in its May research slateResearch & papers · Google DeepMindGoogle DeepMind publishes Co-Scientist and opens Hypothesis GenerationResearch & papers · AgentScout / arXiv trackerRecent arXiv AI-agent papers focus on GUI agents, memory, and multi-agent systemsResearch & papers · arXivCollider-Bench measures whether agents can reproduce LHC analysesResearch & papers · arXivQwen-Scope gives developers sparse-feature tools for Qwen modelsResearch & papers · Hugging Face Papers / arXivSkillRet benchmarks skill retrieval for LLM agentsResearch & papers · arXivForesight Arena proposes an on-chain benchmark for forecasting agentsResearch & papers · arXivMLR-Bench tests whether AI agents can do open-ended ML research