All Research Papers

Browse through 704 research papers

AIDA-ReID: Adaptive Intermediate Domain Adaptation for Generalizable and Source-Free Person Re-Identification

Person re-identification (Re-ID) is challenging due to domain shifts. This paper proposes Adaptive Intermediate Domain Adaptation (AIDA), a framework that treats intermediate-domain learning as a dynamically regulated process. It adaptively controls feature mixing and regularization using feedback signals, and synthesizes diverse intermediate representations with a pseudo-mirror regularization strategy. This approach is demonstrated to be effective across domain generalization and source-free settings, offering significant potential for real-world surveillance and security applications where models need to perform well in unseen environments without access to original training data.

cs.AI
Read More

Privacy-Preserving Federated Learning Framework for Distributed Chemical Process Optimization

This paper proposes a privacy-preserving federated learning framework for distributed chemical process optimization, enabling collaborative model training across multiple geographically separated plants without sharing raw data. The framework significantly improves prediction accuracy across plants and offers a scalable solution for privacy-preserving industrial analytics.

cs.AI
Read More

People-Centred Medical Image Analysis

This paper presents PecMan, a human-AI framework for medical image analysis that optimizes fairness, diagnostic accuracy, and workflow effectiveness. It addresses limited clinical adoption of AI by ensuring fair performance across diverse patient populations and seamless workflow integration through a dynamic gating mechanism.

cs.AI✓ AI Analyzed
Read More

No Digital Content is Safe from Generative AI: Exploring Vulnerabilities in Content Protection

Cybersecurity researchers discovered that simple generative AI tools can easily bypass existing security measures designed to protect digital content from misuse in deepfakes, identity theft, and style mimicry. The study highlights the urgent need for enhanced cybersecurity and trustworthy AI frameworks to counter these vulnerabilities, as current methods offer no foolproof protection against off-the-shelf generative AI models.

cs.AI
Read More

AI for Early Pancreatic Cancer Detection: A Machine Learning Approach with Clinical Data

A new AI model, developed by MIT's Computer Science and Artificial Intelligence Laboratory, trained on routine medical data, can identify patients at high risk of pancreatic cancer up to three years before clinical diagnosis. This breakthrough could enable earlier intervention in a disease with a very low survival rate by recognizing subtle patterns in blood sugar, weight loss, and prescription changes.

cs.AI
Read More

A Novel Computational Framework for Causal Inference: Tree-Based Discretization with ILP-Based Matching

This paper introduces a new computational framework for causal inference combining tree-based discretization and integer linear programming-based matching. It aims to accurately uncover causal relationships from observational data while balancing interpretability and computational efficiency, demonstrating practical advantages over existing methods.

cs.AI
Read More

Towards Lawful Autonomous Driving: Deriving Scenario-Aware Driving Requirements from Traffic Laws and Regulations

This research focuses on systematically deriving scenario-aware driving requirements for autonomous vehicles directly from existing traffic laws and regulations. The aim is to ensure legal compliance and enhance safety for self-driving cars in complex real-world road conditions, providing a critical step towards their widespread and responsible deployment.

cs.AI
Read More

Learning to Rotate: Temporal and Semantic Rotary Encoding for Sequential Modeling

This paper introduces a novel temporal and semantic rotary encoding method designed to improve sequential modeling, offering significant advancements for tasks involving complex time-series data and dynamic systems. Its potential applications range from enhanced natural language processing to more robust robotic control and predictive analytics in various industries.

cs.AI
Read More

FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data

This paper introduces FastOMOP, a foundational architectural framework designed to facilitate reliable and agentic generation of real-world evidence using data harmonized under the OMOP Common Data Model. It aims to accelerate medical research and improve clinical decision-making by providing robust and scalable tools for analyzing diverse healthcare datasets.

cs.AI
Read More

Evaluating whether AI models would sabotage AI safety research

This paper critically examines the potential for advanced AI models to intentionally impede or sabotage efforts in AI safety research. It delves into the risks of misalignment and adversarial behavior from highly capable AI systems, proposing evaluation methods to detect and mitigate such threats, which is vital for the long-term responsible development of AI.

cs.AI
Read More

Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

This work presents a framework for adaptive runtime governance designed to manage and control autonomous AI agents, particularly in scenarios where direct observation of their internal states or decisions is limited. It addresses crucial aspects of safety, reliability, and ethical operation for AI systems deployed in complex and unpredictable real-world environments.

cs.AI
Read More

Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

This research proposes a methodology for developing and validating case-specific rubrics for evaluating clinical AI systems, particularly focusing on the agreement between Large Language Models (LLMs) and clinicians across a large dataset of patient encounters. This work is crucial for the safe and effective deployment of AI in healthcare, ensuring reliable performance in real-world clinical settings.

cs.AI
Read More

Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft

This paper investigates the critical challenge of bridging the gap between AI research discoveries and their practical real-world applications, using Minecraft as a case study for agentic systems. It explores the capabilities of current AI agents in translating theoretical breakthroughs into functional solutions within complex, dynamic environments, offering insights for future AI deployment.

cs.AI
Read More

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity.

This research proposes a novel framework for evaluating mathematical reasoning in large language models (LLMs) that moves beyond rigid symbolic checks, introducing an LLM-as-a-judge paradigm for robust assessment.

cs.AI
Read More

QuantClaw: Precision Where It Matters for OpenClaw.

This paper introduces QuantClaw, a system designed to bring enhanced precision to operations within the OpenClaw framework, focusing on improving accuracy and efficiency in robotic manipulation systems.

cs.AI
Read More

On the Hybrid Nature of ABPMS Process Frames and its Implications on Automated Process Discovery.

This paper explores the hybrid nature of ABPMS process frames and their significant implications for automated process discovery, aiming to enhance the efficiency and accuracy of identifying and understanding business processes.

cs.AI
Read More

Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI.

This paper introduces Defensibility Signals to evaluate rule-governed AI systems, particularly in content moderation, formalizing policy-grounded correctness and offering methods like the Defensibility Index (DI) to assess reasoning stability.

cs.AI
Read More

Architecture of an AI-Based Automated Course of Action Generation System for Military Operations.

This study proposes an architectural design for an AI-based system capable of automatically generating courses of action for military operations, addressing the increasing complexity of modern warfare and the need for rapid decision-making.

cs.AI
Read More

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond.

This paper introduces agentic world modeling, exploring its fundamental principles, current capabilities, and potential future developments, including the establishment of governing laws for AI agents to build and utilize internal representations of the world.

cs.AI
Read More

AgentSearchBench: A Benchmark for AI Agent Search in the Wild.

This paper presents AgentSearchBench, a new benchmark designed to evaluate the performance of AI agents in complex, real-world search scenarios, providing a robust framework for assessing agent capabilities in unconstrained environments.

cs.AI
Read More

CognitiveTwin: Robust Multi-Modal Digital Twins for Predicting Cognitive Decline in Alzheimer's Disease.

This paper introduces CognitiveTwin, a framework for robust multi-modal digital twins designed to predict cognitive decline in Alzheimer's disease, leveraging diverse data sources for early diagnosis and personalized intervention.

cs.AI
Read More

The Silicon Mirror: Dynamic Behavioral Gating for Anti-Sycophancy in LLM Agents

This position paper presents a simulated experimental analysis of AI reliability, focusing on medication decision systems. It introduces dynamic behavioral gating as a mechanism to combat sycophancy in Large Language Model (LLM) agents, aiming to ensure unbiased and safe decision-making in critical applications. The research highlights the importance of robust AI agents in sensitive domains like healthcare, proposing methods to improve their trustworthiness and ethical performance by mitigating undesirable conversational behaviors.

cs.AI
Read More

Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents

This paper proposes a neurosymbolic architecture for domain-grounded AI agents in enterprise systems, integrating ontology-constrained neural reasoning. The approach aims to enhance the trustworthiness and explainability of AI in regulated industries by combining the power of neural networks with symbolic knowledge representation. Empirical evaluations across various industries, including Vietnamese-language domains, demonstrate its potential for creating robust and context-aware AI agents that can operate reliably in complex business environments, ensuring adherence to domain-specific rules and regulations.

cs.AI
Read More

Diagnosing CFG Interpretation in LLMs

This paper investigates the capabilities and limitations of Large Language Models (LLMs) in accurately interpreting Control Flow Graphs (CFGs). It proposes novel diagnostic methodologies to identify common errors and biases in how LLMs process and understand program structures, crucial for enhancing their performance in tasks like code generation, debugging, and vulnerability detection. The research aims to improve the reliability of LLM-powered software development tools, making them more effective and trustworthy for real-world applications in engineering and cybersecurity.

cs.AI
Read More

Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

This paper introduces Uni-SafeBench, a new safety benchmark designed to evaluate unified multimodal large models. It investigates the potential costs associated with the unification of different modalities in large AI models, particularly concerning safety and robustness. The research aims to identify and measure novel safety risks that may emerge from such integrated architectures, providing a critical tool for developers to build more secure and reliable multimodal AI systems for various real-world applications.

cs.AI
Read More

BloClaw: An Omniscient, Multi-Modal Agentic Workspace for Next-Generation Scientific Discovery

This paper introduces BloClaw, an innovative omniscient, multi-modal agentic workspace designed to accelerate next-generation scientific discovery. It leverages advanced AI agents capable of processing and synthesizing information across diverse modalities, from textual literature to experimental data and simulations. The platform aims to provide scientists with a powerful collaborative environment for hypothesis generation, experiment design, and data interpretation, significantly speeding up research cycles and fostering breakthroughs in various scientific fields by enabling a more holistic and intelligent approach to discovery.

cs.AI
Read More

Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

This paper explores agent psychometrics, focusing on predicting task-level performance in agentic coding benchmarks. It delves into methodologies for evaluating the capabilities of AI coding agents beyond simple pass/fail rates, aiming to understand their strengths, weaknesses, and potential for real-world software development. By developing metrics and predictive models for agent performance, the research contributes to building more reliable and efficient AI assistants for programmers, enhancing the overall productivity and quality of software engineering processes.

cs.AI
Read More

Adaptive Parallel Monte Carlo Tree Search for Efficient Test-time Compute Scaling

This paper introduces an adaptive parallel Monte Carlo Tree Search (MCTS) algorithm designed to achieve efficient compute scaling during test-time. It addresses the computational challenges of deploying complex AI decision-making systems in real-world scenarios where resources might be limited or dynamic. The proposed method intelligently distributes computational efforts, enhancing the performance and responsiveness of AI agents in applications such as game AI, autonomous navigation, and strategic planning, making advanced AI more accessible for practical deployment.

cs.AI✓ AI Analyzed
Read More

Towards General-Purpose Representation Learning for Heterogeneous Graphs

This paper investigates the development of general-purpose representation learning techniques specifically designed for heterogeneous graphs. It addresses the challenges of integrating diverse node and edge types in graph neural networks, aiming to create more effective and scalable representations for complex real-world networked data. Such advancements are crucial for applications in social networks, knowledge graphs, and recommendation systems, where data inherently exhibits heterogeneity.

cs.AI
Read More

Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym

This paper introduces Spatial-Gym, a Gymnasium environment that isolates spatial constraint reasoning by testing pathfinding in 2D-grid puzzles as a sequential decision task with optional backtracking. It evaluates AI models in a step-by-step manner, revealing a significant human-model gap in spatial reasoning, and suggesting that current models struggle with global planning when forced into sequential actions. Spatial-Gym provides a framework for diagnosing limitations and improving spatial reasoning through reinforcement learning.

cs.AI
Read More