Variance-Reduced Q-Learning over Static and Time-Varying Networks
Researchers introduced VRDQ, a variance-reduced distributed Q-learning algorithm for multi-agent reinforcement learning over static and dynamic networks.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw Feed or Signals?
Raw Feed is chronological evidence. Signals ranks and interprets material change.
Researchers introduced VRDQ, a variance-reduced distributed Q-learning algorithm for multi-agent reinforcement learning over static and dynamic networks.
CARNet proposes a novel linear-complexity model for multivariate time series forecasting that addresses cross-variate dependencies and periodic patterns.
Research proposes a Decentralized Multi-Agent Swarm (DMAS) architecture using autonomous agents for security in Industrial IoT (IIoT) environments.
Research proposes a method to evaluate the causal impact of ML-assisted decision-making using counterfactual correctness without full RCTs.
Research explores a Spatially-Enhanced Temporal Fusion Transformer for interpretable multi-output prediction in parametric dynamical systems with time-varying inputs.
PCS-UQ introduces a framework for robust uncertainty quantification in high-stakes ML domains, integrating prediction-checks and bootstrapping.
Hopformer, a two-stage Transformer framework, is introduced for multi-variate time series forecasting by separating common trends from specific information.
Research introduces 'adjustment speed' as a safety constraint for reinforcement learning in nonstationary environments, addressing delayed adaptation risks.
A research paper proposes Deep Convolutional Large-Margin $\ell_p$-SVDD for visual anomaly detection, combining deep features with explicit margin-aware boundaries.
Research introduces Hierarchical Online Learning of Multiscale (HOLM) models, combining online latent-cause inference with hierarchical Bayesian models.
Research improves differentially private stochastic gradient descent (DP-SGD) accuracy by correlating privacy noise across iterations using model curvature.
Research introduces Math Education Digital Shadows (MEDS), a dataset to evaluate 14 LLMs' mathematical performance and biases across personifications.
Research introduces SurvDiff, a diffusion model designed for generating synthetic survival data, addressing challenges of incomplete event information.
Research proposes contract-based incentives for federated learning (FL) to prioritize high-quality data contributions during critical early training periods.
Research proposes a convex optimization framework to generate theoretical correlation matrices with graph-based sparsity patterns, improving matrix completion.
Research explores making large text-to-image diffusion models more interpretable and manipulable for creative uses, focusing on interactive explainability.
HiKV proposes a novel algorithm-hardware co-design to compress the KV cache in LLM decoding, tackling memory bottlenecks for long-context models.
Research proposes General Value Functions (GVF) for remaining useful life (RUL) and failure-mode prediction in predictive maintenance.
Research identifies explicit iteration complexity for exact data-driven inverse optimization of Integer Linear Programs using gradient-based methods.
CausalForge is a framework for automated theoretical research in causal inference, designed to address the unreliability of LLM-based reviewers.
Research proposes TRACE-ROUTER, a new routing mechanism for agentic AI applications that optimizes LLM selection based on long-horizon, task-level outcomes.
Research proposes Generalized Gaussian Temporal Difference Error (GGD-TDE) for uncertainty-aware reinforcement learning, addressing non-Gaussian TD residuals.
Research explores meta-learning for speaker-dependent voice fatigue models to improve performance and efficiency over traditional mixed-effect models.
Researchers introduced LiMuon, an optimizer designed for large model training, claiming reduced sample complexity and memory usage compared to prior Muon variants.
Research paper gp2Scale proposes a method for exact Gaussian Processes on up to 10 million data points, improving scalability without approximations.
Research proposes a layer-wise LoRA fine-tuning method using a similarity metric to improve LLM predictive performance efficiently.
DriftXpress introduces an accelerated formulation for 'drifting models,' a new paradigm for one-step generative modeling that reduces inference costs.
Research explores agentic AI for automated, evidence-grounded root cause analysis of industrial anomalies, addressing explainability and data scarcity.
Research paper explores statistical mechanics to better understand extensive-width Bayesian neural networks near interpolation, bridging theory-practice gap.
Research explores safety in In-Context Reinforcement Learning (ICRL), addressing unexamined test-time behavior for real-world deployments.
Researchers propose SEM-DNN, a heteroscedastic neural simultaneous-equation estimator, to learn bidirectional causal interactions from observational data.
Research proposes a theory of indecisions for selective hypothesis testing to minimize abstention rates while maintaining target accuracy in high-risk scenarios.
Research identifies numerical fragility in Transformer models due to low-precision execution, proposing a layer-wise risk estimator and controller.
Re-FORC proposes an adaptive reward prediction method for Chain-of-Thought reasoning to enable early stopping and reduce compute costs by up to 26%.
Research demonstrates Quasi-Monte Carlo (QMC) initialization improves training convergence in meta-reinforcement learning, outperforming orthogonal defaults.
RIS-Kernel introduces a model-agnostic architecture, RIS, reducing LLM self-attention complexity to O(N log N) for long-context inference.
Research paper proposes RED-PIM, a Processing-In-Memory (PIM) architecture to reduce data movement during transformer attention operations, improving efficiency.
Research finds the simple quadratic model can effectively predict optimization dynamics in a 150M parameter LLM, challenging assumptions of neural network complexity.
Research explores multi-horizon latent consistency in video predictors, analyzing how the weighting of multi-step agreement affects prediction error.
Research identifies common metrics for synthetic tabular data generation are blind to inter-column dependencies critical for fraud and risk models.
Research introduces Neural Atom Prevalence (NAP), a Bayesian framework for structured node-level model selection in feedforward neural networks.
Research explores quantum federated learning to enable distributed quantum neural network training without sharing sensitive local data for intelligent services.
Research introduces a parameter-free adaptive sparse attention method using data compression, outperforming fixed patterns and dense attention.
Research proves ReLU networks with two hidden layers can exactly represent the maximum of up to 10 real numbers using rational linear algebra.
Research overviews Bayesian and frequentist simulation-based inference with machine learning for inverse problems and parameter estimation.
Researchers propose Evaluation-as-a-Service (EaaS), a cloud-native microservices architecture for scalable AI monitoring with conformal guarantees.
Research introduces a supervisory runtime stability framework for neural network training to detect and recover from severe destabilizing updates.
Research uses reinforcement learning to replace fixed parameters with state-dependent functions in weather and climate models, improving adaptation to physics.
Research introduces Wasserstein Gradient Flows for scalable and regularized barycenter computation, improving aggregation of probability measures.
Research proposes a new temporal evaluation protocol for personal LLM agents, assessing their evolving memories, skills, and policy states.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion