{"success":true,"data":[{"id":"arxiv:2609.28473","source":"arxiv","source_id":"2609.28473","title":"On the Diffusibility of High-Dimensional Latents","abstract":"Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream effects on generation. Specifically, we show that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient. This motivates using the clean data parameterization ($\\boldsymbol{x}_{0}$-prediction) instead, which focuses learning on the underlying signal manifold. Across experiments with multiple strong-reconstruction encoders, we show that $\\boldsymbol{x}_{0}$-prediction consistently improves text-to-image generation performance.","authors":[{"name":"C"},{"name":"Z"},{"name":"B"},{"name":"Y"},{"name":"X"},{"name":"J"},{"name":"R"},{"name":"Z"},{"name":"A"},{"name":"Y"}],"categories":["cs.CV","cs.LG"],"primary_category":"cs.CV","published_date":"2026-09-23T17:59:56Z","updated_date":"2026-09-23T17:59:56Z","pdf_url":"https://arxiv.org/pdf/2609.28473v1","abs_url":"https://arxiv.org/abs/2609.28473v1","doi":null,"journal_ref":null,"comment":"Accepted to ECCV 2026. Project page: https://cfeng16.github.io/on_the_diffusibility/","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:21","updated_at":"2026-09-24 06:00:21"},{"id":"arxiv:2609.28471","source":"arxiv","source_id":"2609.28471","title":"Contrastive Learning for Authorship Verification","abstract":"Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors of model performance. Based on these considerations, we develop a ModernBERT Bi-Encoder model that achieves 98.4% accuracy on the PAN21 authorship verification task.","authors":[{"name":"P"}],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","published_date":"2026-09-23T17:59:08Z","updated_date":"2026-09-23T17:59:08Z","pdf_url":"https://arxiv.org/pdf/2609.28471v1","abs_url":"https://arxiv.org/abs/2609.28471v1","doi":"10.1007/978-3-032-39150-6_7","journal_ref":"Experimental IR Meets Multilinguality, Multimodality, and Interaction (CLEF 2026), LNCS 17087, pp. 92-102, Springer (2027)","comment":"Published in the proceedings of CLEF 2026. Code: https://github.com/petekirby/contrastive-av","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:22","updated_at":"2026-09-24 06:00:22"},{"id":"arxiv:2609.28470","source":"arxiv","source_id":"2609.28470","title":"StudentBench: AI and human tutoring yield equivalent GRE learning gains","abstract":"Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.","authors":[{"name":"C"},{"name":"I"},{"name":"K"},{"name":"T"},{"name":"A"},{"name":"J"}],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","published_date":"2026-09-23T17:57:45Z","updated_date":"2026-09-23T17:57:45Z","pdf_url":"https://arxiv.org/pdf/2609.28470v1","abs_url":"https://arxiv.org/abs/2609.28470v1","doi":null,"journal_ref":null,"comment":"47 pages, including references and appendices. Project site: https://studentbench.org. GitHub: https://github.com/Handshake-AI-Research/studentbench","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:23","updated_at":"2026-09-24 06:00:23"},{"id":"arxiv:2609.28467","source":"arxiv","source_id":"2609.28467","title":"Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction","abstract":"Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as robotic guide dogs and autonomous mobility scooters. We formulate language-grounded robot group joining: given an observation and a natural-language description of a target group, the robot identifies the relevant group members and predicts socially compliant joining poses. For grounding, we generate structured candidate subsets through recursive spectral partitioning and rank them with a language-conditioned image--geometry model. Given the grounded group, a goal predictor leverages human-formation priors to produce a multimodal energy--orientation map over feasible robot poses. Experiments on conversations, queues, and audiences across varying group sizes, crowd densities, and visual ambiguities show that our method achieves competitive grounding accuracy with sub-second inference and outperforms all baselines in joining-pose prediction. Real-robot experiments further demonstrate group joining in both static and dynamically changing interactions.","authors":[{"name":"Z"},{"name":"Z"},{"name":"G"},{"name":"D"}],"categories":["cs.RO","cs.AI"],"primary_category":"cs.RO","published_date":"2026-09-23T17:55:48Z","updated_date":"2026-09-23T17:55:48Z","pdf_url":"https://arxiv.org/pdf/2609.28467v1","abs_url":"https://arxiv.org/abs/2609.28467v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:26","updated_at":"2026-09-24 06:00:26"},{"id":"arxiv:2609.28466","source":"arxiv","source_id":"2609.28466","title":"The Past Frames the Future: Memory for Autoregressive Video Generation","abstract":"Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities, dynamic states, and intervention-induced causal changes, often leaves the active context long before its relevance diminishes. Overcoming this limitation and maintaining temporal persistence constitutes a fundamental memory problem. We present a systematic and comprehensive review of memory mechanisms in AR video generation. We formulate memory operationally as persistent historical information maintained across outer AR steps, capable of influencing future generation even after the originating evidence is no longer locally accessible. Building upon this unified framework, we organize the literature through five complementary perspectives: (I) Forms, the representational carriers of history; (II) Functions, the specific semantic and physical information requiring preservation; (III) Operations, the lifecycle of writing, reading, updating, managing, and integrating memory; (IV) Learning, the optimization of memory behaviors under closed-loop rollouts; and (V) Evaluation, the paradigms for diagnosing genuine memory capabilities. We conclude by synthesizing open challenges, including composable and resource-aware memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation. By bridging representations, mechanisms, and learning paradigms, this paper establishes a structured foundation for developing reliable, memory-conditioned video generation systems.","authors":[{"name":"H"},{"name":"R"},{"name":"D"},{"name":"W"},{"name":"H"},{"name":"H"},{"name":"S"},{"name":"Z"},{"name":"G"},{"name":"Z"},{"name":"J"},{"name":"Y"},{"name":"R"},{"name":"Y"},{"name":"B"},{"name":"S"},{"name":"Y"},{"name":"S"},{"name":"Y"},{"name":"S"},{"name":"R"},{"name":"N"},{"name":"Y"},{"name":"M"},{"name":"Q"}],"categories":["cs.CV"],"primary_category":"cs.CV","published_date":"2026-09-23T17:54:52Z","updated_date":"2026-09-23T17:54:52Z","pdf_url":"https://arxiv.org/pdf/2609.28466v1","abs_url":"https://arxiv.org/abs/2609.28466v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:26","updated_at":"2026-09-24 06:00:26"},{"id":"arxiv:2609.28459","source":"arxiv","source_id":"2609.28459","title":"Even Sharper Bounds for Transductive Learning and Its Applications","abstract":"We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed-point and confidence terms as the classical inductive local Rademacher-complexity bounds, without the additional logarithmic confidence factor in earlier transductive results. For realizable learning over a binary class of VC dimension $\\dVC$, with training size $m$, test size $u$, and $u\\ge m\\ge\\dVC$, STLC yields $\\cO\\{\\dVC\\log(me/\\dVC)/m\\}$. This matches the standard inductive rate and, when $m\\ge9$, is within a logarithmic factor of the transductive minimax lower bound of order $\\dVC/m$. For transductive kernel learning, STLC gives a spectrum-adaptive excess-risk bound without the multiplicative imbalance factors appearing in the earlier local-complexity bound.","authors":[{"name":"Y"}],"categories":["cs.LG","cs.IT","math.ST"],"primary_category":"cs.LG","published_date":"2026-09-23T17:52:50Z","updated_date":"2026-09-23T17:52:50Z","pdf_url":"https://arxiv.org/pdf/2609.28459v1","abs_url":"https://arxiv.org/abs/2609.28459v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:29","updated_at":"2026-09-24 06:00:29"},{"id":"arxiv:2609.28449","source":"arxiv","source_id":"2609.28449","title":"Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark","abstract":"Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. Finally, we show that the oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.","authors":[{"name":"H"},{"name":"M"},{"name":"M"},{"name":"H"},{"name":"T"},{"name":"H"}],"categories":["cs.SE","cs.AI","cs.CL"],"primary_category":"cs.SE","published_date":"2026-09-23T17:46:50Z","updated_date":"2026-09-23T17:46:50Z","pdf_url":"https://arxiv.org/pdf/2609.28449v1","abs_url":"https://arxiv.org/abs/2609.28449v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:29","updated_at":"2026-09-24 06:00:29"},{"id":"arxiv:2609.28448","source":"arxiv","source_id":"2609.28448","title":"Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality","abstract":"We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similarity-based attention selects nearby representations, while the negative value map drives tokens away from the selected field. This feedback can continually reorganize both the representation geometry and the attention network. For $d=2$, the tokens lie on a circle, where the regular polygon is an exact fixed point. As the attention feedback strength $γ$ is increased, the polygon loses stability through a flip bifurcation, giving rise to period-two motion, chaos, and cluster-exchange or cluster-flip states. Despite this temporal complexity, attention remains diffuse as $N\\to\\infty$ at finite fixed softmax sharpness $β$. Attention condensation instead emerges in the scaling regime $β\\sim N^2$. In the hard-routing limit, repulsive updates amplify local perturbations and routing-partner switches transmit them ballistically, producing an emergent butterfly cone in representation space. High-dimensional geometry provides a distinct route to localization. For $d=N\\to\\infty$, simulations from Gaussian initial conditions provide evidence for a condensation transition at $β=O(1)$, driven by dynamically generated finite overlap gaps. Depending on $γ$, the resulting phases include diffuse simplex-like states, consensus flips, condensed active routing with signatures of chaos, and fragmented cluster flips. These results establish temporal activity, attention condensation, and geometric clustering as distinct collective phenomena, and show that sparse attention can sustain persistent dynamics rather than freeze it.","authors":[{"name":"Q"},{"name":"Z"},{"name":"X"}],"categories":["cond-mat.dis-nn","cond-mat.stat-mech","cs.LG"],"primary_category":"cond-mat.dis-nn","published_date":"2026-09-23T17:46:38Z","updated_date":"2026-09-23T17:46:38Z","pdf_url":"https://arxiv.org/pdf/2609.28448v1","abs_url":"https://arxiv.org/abs/2609.28448v1","doi":null,"journal_ref":null,"comment":"54 pages, 21 figures, including appendices","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:30","updated_at":"2026-09-24 06:00:30"},{"id":"arxiv:2609.28442","source":"arxiv","source_id":"2609.28442","title":"Order-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning","abstract":"Reordering a set of mathematical rules without changing its meaning should preserve the correct answer, but must a model's internal representations stay invariant too? We investigate this question using synthetic multi-step function-composition problems, each presented under multiple rule orderings with the same correct answer. We measure accuracy and permutation signal-to-noise ratio (SNR), which quantifies how distinctly ordering patterns are represented relative to variation across problem instances. Across 16 language models ranging from 1B to 8B parameters, we find a pattern: models that solve reordered problems more accurately represent different rule orderings more distinctly. Layer-averaged permutation SNR is positively rank-correlated with accuracy in every synthetic setting we evaluate, with Spearman correlations reaching 0.86. These findings highlight a distinction between answer invariance and representation invariance: successful mathematical rule composition can accompany distinct internal representations between equivalent rule orderings. This motivates distinguishing answer invariance from representation invariance, and offers a representational perspective on mathematical reasoning beyond answer accuracy alone.","authors":[{"name":"Z"}],"categories":["cs.LG","cs.AI","cs.CL","cs.SC"],"primary_category":"cs.LG","published_date":"2026-09-23T17:40:57Z","updated_date":"2026-09-23T17:40:57Z","pdf_url":"https://arxiv.org/pdf/2609.28442v1","abs_url":"https://arxiv.org/abs/2609.28442v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:30","updated_at":"2026-09-24 06:00:30"},{"id":"arxiv:2609.28439","source":"arxiv","source_id":"2609.28439","title":"HaRP: High Dynamic Range Photosequencing through Dual Reversed Shutter Scanning","abstract":"The adoption of CMOS sensors in mobile photography is frequently compromised by the rolling shutter (RS) effect, which introduces geometric distortions and motion artifacts. Particularly, recent rolling shutter with global reset (RSGR) mode, while mitigating some RS issues, also incurs major limitations, including reduced capture speed and compressed dynamic range. To address these problems, we propose a novel dual reversed scanning setup utilizing both RSGR and inverted RSGR views. This solution not only handles the inherent flaws of RSGR by synchronizing complementary exposures to balance the dynamic range across the frames but also introduces an effective method for HDR photosequencing under highly dynamic scenes. Our proposed network first accommodates row-wise complementarity and manages visual shifts by row-adaptive feature alignment. Subsequently, the hallucination module, built upon a correlation-guided mixattention block, integrates the mutually reinforced features to recover missing details. In addition, we construct a coaxial imaging system to collect a real-world dataset, enabling robust training and evaluation beyond numerical simulation. Experimental results demonstrate the twofold benefits of our solution in mitigating RSGR limitations and advancing HDR reconstruction techniques.","authors":[{"name":"X"},{"name":"G"},{"name":"J"},{"name":"Z"},{"name":"Y"}],"categories":["cs.CV"],"primary_category":"cs.CV","published_date":"2026-09-23T17:38:30Z","updated_date":"2026-09-23T17:38:30Z","pdf_url":"https://arxiv.org/pdf/2609.28439v1","abs_url":"https://arxiv.org/abs/2609.28439v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:31","updated_at":"2026-09-24 06:00:31"},{"id":"arxiv:2609.28438","source":"arxiv","source_id":"2609.28438","title":"Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections","abstract":"We study minimal-norm interpolation and $\\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-layer biases are included in the parameter norm. When biases are unpenalized, the minimal-norm interpolators are exactly the continuous piecewise-affine functions that hug every label switch and have kinks of the appropriate convexity. When biases are penalized, the minimizer is unique in function space, has exactly one kink in each intermediate same-label segment, and is therefore a sparsest positive-margin classifier. We further show that adding a free affine skip connection leaves these function-space solutions unchanged but fundamentally improves the parameter-space landscape: every KKT point of the constrained problem becomes globally optimal, whereas suboptimal KKT points can occur without the skip connection. We establish analogous global-optimality and geometric results for sufficiently weak $\\ell_2$-regularization of the logistic loss. In the unpenalized-bias case, we identify an additional sparsity-like restriction, implying that most minimal-norm interpolators cannot arise as small-regularization limits of margin-normalized logistic-loss minimizers. Numerical experiments across varying dataset complexity and network width support the predicted landscape and sparsity phenomena.","authors":[{"name":"K"},{"name":"B"},{"name":"A"},{"name":"E"},{"name":"P"},{"name":"M"},{"name":"R"}],"categories":["cs.LG"],"primary_category":"cs.LG","published_date":"2026-09-23T17:37:26Z","updated_date":"2026-09-23T17:37:26Z","pdf_url":"https://arxiv.org/pdf/2609.28438v1","abs_url":"https://arxiv.org/abs/2609.28438v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:32","updated_at":"2026-09-24 06:00:32"},{"id":"arxiv:2609.28437","source":"arxiv","source_id":"2609.28437","title":"MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos","abstract":"Online information is increasingly consumed in video format. Much of this comes in the form of *raw video*: continuous footage taken on a cell phone, with a hand-held camera, or via CCTV, which is then directly uploaded to social media platforms and content sharing services. Whereas professional or even amateur-edited footage tends to feature scripted speech, chyrons, graphics, and metadata that help contextualize its subject matter, raw video typically contains none of these things, making it a much more challenging medium for information retrieval and machine understanding. To facilitate progress in this domain, we release MultiVENT-Raw, a multilingual collection of nearly 120,000 primarily raw videos (over 5,300 total hours), paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos. MultiVENT-Raw supports both a retrieval task---to identify videos in the collection relevant to a query event---and a generation task---to summarize event-related videos into a coherent report for a target user. We benchmark strong baselines on MultiVENT-Raw, showing both tasks to be challenging even for some of the latest multimodal models.","authors":[{"name":"R"},{"name":"D"},{"name":"A"},{"name":"C"},{"name":"D"},{"name":"H"},{"name":"R"},{"name":"M"},{"name":"K"},{"name":"E"},{"name":"B"},{"name":"A"},{"name":"A"},{"name":"W"}],"categories":["cs.CV","cs.IR"],"primary_category":"cs.CV","published_date":"2026-09-23T17:35:40Z","updated_date":"2026-09-23T17:35:40Z","pdf_url":"https://arxiv.org/pdf/2609.28437v1","abs_url":"https://arxiv.org/abs/2609.28437v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:32","updated_at":"2026-09-24 06:00:32"},{"id":"arxiv:2609.28434","source":"arxiv","source_id":"2609.28434","title":"Predicting the Progression of Adolescent Idiopathic Scoliosis","abstract":"Adolescent Idiopathic Scoliosis is defined as a lateral curvature of the spine that develops during adolescence, without known cause. The condition can result in significant pain and disability, and often progresses rapidly during adolescence. The objective of this paper is to predict the progression of the condition in a temporal sequence from ages 9 to 24, as measured from a sequence of Dual X-ray Absorptiometry (DXA) scans. To this end, we train a transformer model that takes in the curve of the spine to predict curve progression. The model is trained using a large-scale synthetic dataset of spine curves and their time series, covering different curve types and different progression patterns. We show that the model is able to generalise from synthetic to real data by evaluating it on a dataset of real DXA scans covering multiple time points. We find that fine-tuning the model on real data gives a significant boost to performance. The model is able to accurately predict spine curve progression in both scoliosis and normal cases.","authors":[{"name":"O"},{"name":"A"},{"name":"A"}],"categories":["cs.CV","eess.IV"],"primary_category":"cs.CV","published_date":"2026-09-23T17:33:31Z","updated_date":"2026-09-23T17:33:31Z","pdf_url":"https://arxiv.org/pdf/2609.28434v1","abs_url":"https://arxiv.org/abs/2609.28434v1","doi":null,"journal_ref":null,"comment":"Published in MICCAI ShapeMI 2026 Workshop","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:33","updated_at":"2026-09-24 06:00:33"},{"id":"arxiv:2609.28431","source":"arxiv","source_id":"2609.28431","title":"LiMA: Bridging Long-term Imagination to Real-time Dexterous Manipulation via Asynchronous Diffusion","abstract":"Dexterous manipulation demands long-term foresight and rapid reactive control. Vision-Language-Action (VLA) models, while proficient in high-level reasoning, often lack a fine-grained understanding of physical dynamics and spatial perception. Conversely, World-Action Models (WAMs) typically suffer from high inference latency due to iterative generation. These deficiencies result in a critical temporal misalignment where the model's intent fails to adapt to rapid physical contact changes. To overcome this fundamental bottleneck, we propose LiMA, an asynchronous dual-system generative framework that systematically decouples intent planning from reactive execution. LiMA organizes computation into a multi-scale hierarchy: a slow system handles sparse long-horizon spatiotemporal intent generation, while a fast system focuses on dense high-frequency motion refinement. To align sparse intent predictions with dense action trajectories, we introduce a Latent Schrödinger Bridge Coupling mechanism that formulates refinement as an entropy-regularized probabilistic transport process. LiMA reduces inference latency by 45.8% compared with Cosmos-Policy via asynchronous decoupling. Evaluated across six bimanual dexterous manipulation tasks spanning multiple horizons, LiMA achieves an overall success rate of 70.8% and an average subtask success rate of 78.9%, while maintaining performance in unseen scenarios. The project website is available at https://ccdcs.github.io/LiMA_repo/","authors":[{"name":"N"},{"name":"Y"},{"name":"J"},{"name":"Q"},{"name":"G"},{"name":"P"},{"name":"Z"},{"name":"S"}],"categories":["cs.RO"],"primary_category":"cs.RO","published_date":"2026-09-23T17:31:52Z","updated_date":"2026-09-23T17:31:52Z","pdf_url":"https://arxiv.org/pdf/2609.28431v1","abs_url":"https://arxiv.org/abs/2609.28431v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:34","updated_at":"2026-09-24 06:00:34"},{"id":"arxiv:2609.28430","source":"arxiv","source_id":"2609.28430","title":"Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms","abstract":"This work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential low-rank adaptation (LoRA) protocol for cross-scale transfer: a Qwen3 backbone with a bounded regression head is first fine-tuned on the English DAIC-WOZ dataset (189 avatar-mediated sessions, PHQ-8), and the adapter then initializes fine-tuning on the Chinese PDCH dataset (100 real clinical consultations, HAMD-17), where a reinitialised, scale-specific head predicts the clinician-assigned score. All configurations use patient-level stratified 5-fold, 2-repeat cross-validation. On the data-scarce HAMD-17 target, the sequential protocol attains the best point-estimate MAE , RMSE, and macro-$F_1$ on both 0.6B and 1.7B backbones, outperforming target-only training and non-LLM baselines---4.96/6.59/0.36 with Qwen3-0.6B and 4.38/5.62/0.46 with Qwen3-1.7B. Ablations suggest that correctly aligned source supervision gives the best point estimates (unsupervised exposure and shuffled-label controls also show partial gains), that native-Chinese target input outperforms machine-translated English input, and that the reversed order yields no clear gain within run-to-run variance. The study is an exploratory, single-site internal evaluation: it does not establish screening or diagnostic utility, nor separately identify the contribution of the scale, language, or paradigm shifts. To our knowledge, no prior study evaluates this specific DAIC-WOZ-to-PDCH sequential transfer setting.","authors":[{"name":"W"},{"name":"S"},{"name":"S"}],"categories":["cs.CL"],"primary_category":"cs.CL","published_date":"2026-09-23T17:30:07Z","updated_date":"2026-09-23T17:30:07Z","pdf_url":"https://arxiv.org/pdf/2609.28430v1","abs_url":"https://arxiv.org/abs/2609.28430v1","doi":null,"journal_ref":null,"comment":"preprint to ICASSP 2027","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:34","updated_at":"2026-09-24 06:00:34"},{"id":"arxiv:2609.28429","source":"arxiv","source_id":"2609.28429","title":"Watch, Recall, Act: Always-On Robots in Concurrent Embodied Streams","abstract":"An always-on robot faces an endless stream that never resets: instructions arrive and lapse, the scene changes, and its own past actions reshape what it must reason about. Today's action models are built for the opposite: a fixed instruction, no mid-task intervention, single-step reasoning. In an open-ended world a robot must watch a live stream for far-future cues, recall its own far-past actions, and act on them under dual-arm concurrency. We present ARMS (Always-on Robot in Multi-modal Streams), a deliberately simple streaming policy: a single pretrained $π$0.5 backbone augmented by three lightweight modules that turn live perception, embodied states, and the robot's own past actions into context the backbone reads before it acts. The modules update this context asynchronously, so watching and recalling never block acting and the two arms act at once. Rather than inventing new mechanisms, ARMS integrates these learned context providers with an agent-causal self-history that logs which arm did what, and when. To supervise them without extra annotation, we build ARMS Dataset, whose staged construction script itself labels every module from real dual-arm teleoperation. Trained on it, ARMS reaches 45% on the combined task against 28% for the strongest of our four main baselines, and ablations confirm the memory module, the embodied-state head, and asynchronous concurrency are each necessary.","authors":[{"name":"D"},{"name":"P"},{"name":"C"},{"name":"J"},{"name":"X"},{"name":"X"},{"name":"X"}],"categories":["cs.RO"],"primary_category":"cs.RO","published_date":"2026-09-23T17:29:40Z","updated_date":"2026-09-23T17:29:40Z","pdf_url":"https://arxiv.org/pdf/2609.28429v1","abs_url":"https://arxiv.org/abs/2609.28429v1","doi":null,"journal_ref":null,"comment":"11 pages, 3 figures. Accepted to the 10th Conference on Robot Learning (CoRL 2026)","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:35","updated_at":"2026-09-24 06:00:35"},{"id":"arxiv:2609.28427","source":"arxiv","source_id":"2609.28427","title":"Context-Continuous Preference Learning for Exoskeleton Personalization","abstract":"Personalizing exoskeleton assistance across operating conditions is constrained by the time and physical effort required to collect user feedback. We examined whether a user's preference landscape varies smoothly across operating conditions and when this continuity supports learning from limited feedback. We propose Context-Continuous Preference Learning (CCPL), a Gaussian-process preference model that shares observations across nearby contexts while retaining context-specific utility estimates. We evaluated CCPL through simulations and retrospective analyses of ankle and elbow exoskeleton preference data from nine healthy adults. In simulations, CCPL improved reconstruction and preference-based Bayesian optimization relative to independent learning when preferences varied smoothly, but showed negative transfer when continuity was weak. In both human studies, full-data reference landscapes estimated separately for each participant and context tended to be more similar between nearby operating conditions. With five exposures per context, CCPL increased mean reconstruction correlation with these references from 0.644 to 0.720 for ankle assistance and from 0.476 to 0.526 for elbow assistance relative to independent learning. The five-exposure budget was approximately 37% lower for ankle and 17% lower for elbow than the estimated independent-learning budgets needed to match these correlations. CCPL also improved held-out response prediction relative to independent learning, while benefits over pooled learning varied. These findings support context continuity as a basis for sharing preference observations under limited feedback, although benefits for online personalization in humans remain to be established.","authors":[{"name":"S"},{"name":"S"},{"name":"D"}],"categories":["cs.LG","cs.RO"],"primary_category":"cs.LG","published_date":"2026-09-23T17:28:08Z","updated_date":"2026-09-23T17:28:08Z","pdf_url":"https://arxiv.org/pdf/2609.28427v1","abs_url":"https://arxiv.org/abs/2609.28427v1","doi":null,"journal_ref":null,"comment":"23 pages, 11 figures, including supplementary materials","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:36","updated_at":"2026-09-24 06:00:36"},{"id":"arxiv:2609.28425","source":"arxiv","source_id":"2609.28425","title":"Repairability of Inexact Solvers in Recursive State Estimation with Machine Learning","abstract":"Recursive state estimation often executes approximate numerical solutions inside a feedback loop, where highly accurate local steps do not guarantee better overall results. For a fixed linear Kalman model, we characterize when a correction within a prescribed subspace and norm budget can meet a local admissibility tolerance, and how the defects actually executed affect the finite-horizon covariance response. Centering each defect on the exact gain for the implemented covariance separates current solve error from inherited gain drift. Expanding the exact residual-drift identity reveals opposing quartic contributions beyond the quadratic response: innovation-covariance inflation enters positively, while local-gain reoptimization enters subtractively. Under matched initialization, an absolute sixth-order remainder bound, uniform over bounded defect sequences at fixed horizon, gives sufficient conditions for quadratic under- or overprediction. Machine learning proposes bounded corrections, while a learner-independent residual certificate and verified fallback govern execution of classical and quantum candidates without changing the reference estimator. In a power-grid tolerance study, learned correction lowers the minimum conjugate-gradient iteration count for deployment without fallback relative to uncorrected solves under the same residual certificate. Gains reconstructed from a variational quantum linear solver and from an annealing-based binary encoding, with small-scale terminal measurements on superconducting hardware and sampling on a quantum annealer, are executed through the same interface. By linking local repairability to nonlinear error propagation, the framework evaluates approximate solvers and learned corrections through independent certification and finite-horizon response, providing a practical basis for studying hybrid quantum--classical computation.","authors":[{"name":"Y"},{"name":"D"},{"name":"O"},{"name":"P"},{"name":"Z"},{"name":"M"},{"name":"I"},{"name":"F"},{"name":"B"},{"name":"C"},{"name":"K"}],"categories":["quant-ph","cs.LG"],"primary_category":"quant-ph","published_date":"2026-09-23T17:27:07Z","updated_date":"2026-09-23T17:27:07Z","pdf_url":"https://arxiv.org/pdf/2609.28425v1","abs_url":"https://arxiv.org/abs/2609.28425v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:36","updated_at":"2026-09-24 06:00:36"},{"id":"arxiv:2609.28424","source":"arxiv","source_id":"2609.28424","title":"The Skin-Restricted Reinhard Transform:Uniqueness under a Lightness-Preserving Constraint","abstract":"Catalog skin recolouring has to change pigment and leave shading alone. The classical Reinhard map does not make that split: it rescales lightness by the ratio of standard deviations, and a flat reference swatch therefore flattens the limb. This paper formalises the correction used in our pipeline, the skin-restricted Reinhard transform. It is the diagonal affine map in CIE Lab that translates lightness, matches the chromatic mean, and clamps the chromatic gain to [0.72, 1.18], with moments taken on the central 84% of each channel. A diagonal affine map has six real parameters. The shading constraint forces the lightness gain to +1 and the lightness shift to the difference of means; one-dimensional quadratic optimal transport on each chromatic axis, followed by Euclidean projection onto the gain interval, fixes the other four. Inside that family the four conditions determine every parameter. The content of the result is the forced lightness gain; it is not a uniqueness claim outside the diagonal affine class. For Gaussian marginals the chromatic step is not merely the best affine map: it is the unrestricted Wasserstein-2 map. The same formulae with trimmed moments remain optimal because a positive affine image commutes with quantile trimming. On hands, arms, legs, and feet of nine photographs and three reference tones, the map keeps the lightness contrast ratio at 0.974 +/- 0.029 with chromatic error 0.77 CIE Lab units. Reinhard matching, the linear Monge map, and histogram matching reach a smaller chromatic error only by cutting lightness contrast to about half.","authors":[{"name":"V"}],"categories":["cs.CV"],"primary_category":"cs.CV","published_date":"2026-09-23T17:27:01Z","updated_date":"2026-09-23T17:27:01Z","pdf_url":"https://arxiv.org/pdf/2609.28424v1","abs_url":"https://arxiv.org/abs/2609.28424v1","doi":null,"journal_ref":null,"comment":"Code: https://github.com/vijeshkpaei/skin-restricted-reinhard-transform","pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:37","updated_at":"2026-09-24 06:00:37"},{"id":"arxiv:2609.28416","source":"arxiv","source_id":"2609.28416","title":"Agent-Editing World Model: Rethinking World Modeling for LLM Agents","abstract":"Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \\emph{task-state contamination}, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the \\textbf{Agent-Editing World Model (AEWM)}, which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines \\textbf{Action Judge} to distinguish \\textsc{Critical}, \\textsc{Exploratory}, and \\textsc{Noisy} decisions with \\textbf{State Revision} to edit noisy reasoning--action continuations from the same observed history. \\textbf{EditAct} integrates these capabilities with real execution, directly changing the state underlying subsequent decisions rather than merely providing critiques. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5\\% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline. Furthermore, rejection sampling fine-tuning on verified EditAct trajectories, termed \\textbf{AEWM-RFT}, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.","authors":[{"name":"S"},{"name":"G"},{"name":"F"},{"name":"J"},{"name":"H"},{"name":"J"},{"name":"W"},{"name":"H"},{"name":"J"}],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","published_date":"2026-09-23T17:18:26Z","updated_date":"2026-09-23T17:18:26Z","pdf_url":"https://arxiv.org/pdf/2609.28416v1","abs_url":"https://arxiv.org/abs/2609.28416v1","doi":null,"journal_ref":null,"comment":null,"pdf_downloaded":false,"pdf_path":null,"trending_score":0,"citation_count":0,"github_stars":0,"created_at":"2026-09-24 06:00:37","updated_at":"2026-09-24 06:00:37"}],"meta":{"page":1,"per_page":20,"total":7400,"total_pages":370}}