subscribe to arXiv mailings

15M Multimodal Facial Image-Text Dataset

Authors: Dawei Dai, YuTang Li, YingGe Liu, Mingming Jia, Zhang YuanHui, Guoyin Wang

Abstract: Currently, image-text-driven multi-modal deep learning models have demonstrated their outstanding potential in many fields. In practice, tasks centered around facial images have broad application prospects. This paper presents \textbf{FaceCaption-15M}, a large-scale, diverse, and high-quality dataset of facial images accompanied by their natural language descriptions (facial image-to-text). This d… ▽ More Currently, image-text-driven multi-modal deep learning models have demonstrated their outstanding potential in many fields. In practice, tasks centered around facial images have broad application prospects. This paper presents \textbf{FaceCaption-15M}, a large-scale, diverse, and high-quality dataset of facial images accompanied by their natural language descriptions (facial image-to-text). This dataset aims to facilitate a study on face-centered tasks. FaceCaption-15M comprises over 15 million pairs of facial images and their corresponding natural language descriptions of facial features, making it the largest facial image-caption dataset to date. We conducted a comprehensive analysis of image quality, text naturalness, text complexity, and text-image relevance to demonstrate the superiority of FaceCaption-15M. To validate the effectiveness of FaceCaption-15M, we first trained a facial language-image pre-training model (FLIP, similar to CLIP) to align facial image with its corresponding captions in feature space. Subsequently, using both image and text encoders and fine-tuning only the linear layer, our FLIP-based models achieved state-of-the-art results on two challenging face-centered tasks. The purpose is to promote research in the field of face-related tasks through the availability of the proposed FaceCaption-15M dataset. All data, codes, and models are publicly available. https://huggingface.co/datasets/OpenFace-CQUPT/FaceCaption-15M △ Less

Submitted 11 July, 2024; v1 submitted 11 July, 2024; originally announced July 2024.

Comments: 15 pages, 8 figures

arXiv:2407.08132 [pdf, other]

DMM: Disparity-guided Multispectral Mamba for Oriented Object Detection in Remote Sensing

Authors: Minghang Zhou, Tianyu Li, Chaofan Qiao, Dongyu Xie, Guoqing Wang, Ningjuan Ruan, Lin Mei, Yang Yang

Abstract: Multispectral oriented object detection faces challenges due to both inter-modal and intra-modal discrepancies. Recent studies often rely on transformer-based models to address these issues and achieve cross-modal fusion detection. However, the quadratic computational complexity of transformers limits their performance. Inspired by the efficiency and lower complexity of Mamba in long sequence task… ▽ More Multispectral oriented object detection faces challenges due to both inter-modal and intra-modal discrepancies. Recent studies often rely on transformer-based models to address these issues and achieve cross-modal fusion detection. However, the quadratic computational complexity of transformers limits their performance. Inspired by the efficiency and lower complexity of Mamba in long sequence tasks, we propose Disparity-guided Multispectral Mamba (DMM), a multispectral oriented object detection framework comprised of a Disparity-guided Cross-modal Fusion Mamba (DCFM) module, a Multi-scale Target-aware Attention (MTA) module, and a Target-Prior Aware (TPA) auxiliary task. The DCFM module leverages disparity information between modalities to adaptively merge features from RGB and IR images, mitigating inter-modal conflicts. The MTA module aims to enhance feature representation by focusing on relevant target regions within the RGB modality, addressing intra-modal variations. The TPA auxiliary task utilizes single-modal labels to guide the optimization of the MTA module, ensuring it focuses on targets and their local context. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate the effectiveness of our method, which outperforms state-of-the-art methods while maintaining computational efficiency. Code will be available at https://github.com/Another-0/DMM. △ Less

Submitted 10 July, 2024; originally announced July 2024.

Comments: 12 pages, 9 figures

arXiv:2407.07025 [pdf, other]

Quantum Frequency Mixing using an NV Diamond Microscope

Authors: Samuel J. Karlson, Pauli Kehayias, Jennifer M. Schloss, Andrew C. Maccabe, David F. Phillips, Guoqing Wang, Paola Cappellaro, Danielle A. Braje

Abstract: Wide-field magnetic microscopy using nitrogen-vacancy (NV) centers in diamond can yield high-quality magnetic images of DC and AC magnetic fields. The unique combination of micron-scale spatial resolution of scalar or vector fields at room temperature and parallel camera readout make this an appealing technique for applications in biology, geology, condensed-matter physics, and electronics. Howeve… ▽ More Wide-field magnetic microscopy using nitrogen-vacancy (NV) centers in diamond can yield high-quality magnetic images of DC and AC magnetic fields. The unique combination of micron-scale spatial resolution of scalar or vector fields at room temperature and parallel camera readout make this an appealing technique for applications in biology, geology, condensed-matter physics, and electronics. However, while NV magnetic microscopy has achieved great success in these areas, historically the accessible frequency range has been limited. In this paper, we overcome this limitation by implementing the recently developed technique of quantum frequency mixing. With this approach, we generate wide-field magnetic images of test structures driven by alternating currents up to 70 MHz, well outside the reach of DC and Rabi magnetometry methods. With further improvements, this approach could find utility in hyperspectral imaging for electronics power spectrum analysis, electronics diagnostics and troubleshooting, and quantum computing hardware validation. △ Less

Submitted 9 July, 2024; originally announced July 2024.

Comments: 8 pages main text, 3 pages supplement

arXiv:2407.06988 [pdf, other]

Limiting Over-Smoothing and Over-Squashing of Graph Message Passing by Deep Scattering Transforms

Authors: Yuanhong Jiang, Dongmian Zou, Xiaoqun Zhang, Yu Guang Wang

Abstract: Graph neural networks (GNNs) have become pivotal tools for processing graph-structured data, leveraging the message passing scheme as their core mechanism. However, traditional GNNs often grapple with issues such as instability, over-smoothing, and over-squashing, which can degrade performance and create a trade-off dilemma. In this paper, we introduce a discriminatively trained, multi-layer Deep… ▽ More Graph neural networks (GNNs) have become pivotal tools for processing graph-structured data, leveraging the message passing scheme as their core mechanism. However, traditional GNNs often grapple with issues such as instability, over-smoothing, and over-squashing, which can degrade performance and create a trade-off dilemma. In this paper, we introduce a discriminatively trained, multi-layer Deep Scattering Message Passing (DSMP) neural network designed to overcome these challenges. By harnessing spectral transformation, the DSMP model aggregates neighboring nodes with global information, thereby enhancing the precision and accuracy of graph signal processing. We provide theoretical proofs demonstrating the DSMP's effectiveness in mitigating these issues under specific conditions. Additionally, we support our claims with empirical evidence and thorough frequency analysis, showcasing the DSMP's superior ability to address instability, over-smoothing, and over-squashing. △ Less

Submitted 9 July, 2024; originally announced July 2024.

Comments: 35 pages, 6 figures

arXiv:2407.06964 [pdf, other]

Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach

Authors: Taolin Zhang, Jiawang Bai, Zhihe Lu, Dongze Lian, Genping Wang, Xinchao Wang, Shu-Tao Xia

Abstract: Recent works on parameter-efficient transfer learning (PETL) show the potential to adapt a pre-trained Vision Transformer to downstream recognition tasks with only a few learnable parameters. However, since they usually insert new structures into the pre-trained model, entire intermediate features of that model are changed and thus need to be stored to be involved in back-propagation, resulting in… ▽ More Recent works on parameter-efficient transfer learning (PETL) show the potential to adapt a pre-trained Vision Transformer to downstream recognition tasks with only a few learnable parameters. However, since they usually insert new structures into the pre-trained model, entire intermediate features of that model are changed and thus need to be stored to be involved in back-propagation, resulting in memory-heavy training. We solve this problem from a novel disentangled perspective, i.e., dividing PETL into two aspects: task-specific learning and pre-trained knowledge utilization. Specifically, we synthesize the task-specific query with a learnable and lightweight module, which is independent of the pre-trained model. The synthesized query equipped with task-specific knowledge serves to extract the useful features for downstream tasks from the intermediate representations of the pre-trained model in a query-only manner. Built upon these features, a customized classification head is proposed to make the prediction for the input sample. lightweight architecture and avoids the use of heavy intermediate features for running gradient descent, it demonstrates limited memory usage in training. Extensive experiments manifest that our method achieves state-of-the-art performance under memory constraints, showcasing its applicability in real-world situations. △ Less

Submitted 9 July, 2024; originally announced July 2024.

Comments: ECCV2024

arXiv:2407.06752 [pdf, ps, other]

On Arnold-type stability theorems for the Euler equation on a sphere

Authors: Daomin Cao, Guodong Wang

Abstract: In this paper, we establish three Arnold-type stability theorems for steady or rotating solutions of the incompressible Euler equation on a sphere. Specifically, we prove that if the stream function of a flow solves a semilinear elliptic equation with a monotone nonlinearity, then, under appropriate conditions, the flow is stable or orbitally stable in the Lyapunov sense. In particular, our theore… ▽ More In this paper, we establish three Arnold-type stability theorems for steady or rotating solutions of the incompressible Euler equation on a sphere. Specifically, we prove that if the stream function of a flow solves a semilinear elliptic equation with a monotone nonlinearity, then, under appropriate conditions, the flow is stable or orbitally stable in the Lyapunov sense. In particular, our theorems apply to degree-2 Rossby-Haurwitz waves. These results are achieved via a variational approach, with the key ingredient being to show that the flows under consideration satisfy the conditions of two Burton-type stability criteria which are established in this paper. As byproducts, we obtain some sharp rigidity results for solutions of semilinear elliptic equations on a sphere. △ Less

Submitted 9 July, 2024; originally announced July 2024.

Comments: 27 pages

arXiv:2407.06503 [pdf, other]

Preference-Guided Reinforcement Learning for Efficient Exploration

Authors: Guojian Wang, Faguo Wu, Xiao Zhang, Tianyuan Chen, Xuyang Chen, Lin Zhao

Abstract: In this paper, we investigate preference-based reinforcement learning (PbRL) that allows reinforcement learning (RL) agents to learn from human feedback. This is particularly valuable when defining a fine-grain reward function is not feasible. However, this approach is inefficient and impractical for promoting deep exploration in hard-exploration tasks with long horizons and sparse rewards. To tac… ▽ More In this paper, we investigate preference-based reinforcement learning (PbRL) that allows reinforcement learning (RL) agents to learn from human feedback. This is particularly valuable when defining a fine-grain reward function is not feasible. However, this approach is inefficient and impractical for promoting deep exploration in hard-exploration tasks with long horizons and sparse rewards. To tackle this issue, we introduce LOPE: Learning Online with trajectory Preference guidancE, an end-to-end preference-guided RL framework that enhances exploration efficiency in hard-exploration tasks. Our intuition is that LOPE directly adjusts the focus of online exploration by considering human feedback as guidance, avoiding learning a separate reward model from preferences. Specifically, LOPE includes a two-step sequential policy optimization process consisting of trust-region-based policy improvement and preference guidance steps. We reformulate preference guidance as a novel trajectory-wise state marginal matching problem that minimizes the maximum mean discrepancy distance between the preferred trajectories and the learned policy. Furthermore, we provide a theoretical analysis to characterize the performance improvement bound and evaluate the LOPE's effectiveness. When assessed in various challenging hard-exploration environments, LOPE outperforms several state-of-the-art methods regarding convergence rate and overall performance. The code used in this study is available at \url{https://github.com/buaawgj/LOPE}. △ Less

Submitted 8 July, 2024; originally announced July 2024.

Comments: 13 pages, 17 figures

arXiv:2407.05810 [pdf, other]

Integrating AI in College Education: Positive yet Mixed Experiences with ChatGPT

Authors: Xinrui Song, Jiajin Zhang, Pingkun Yan, Juergen Hahn, Uwe Kruger, Hisham Mohamed, Ge Wang

Abstract: The integration of artificial intelligence (AI) chatbots into higher education marks a shift towards a new generation of pedagogical tools, mirroring the arrival of milestones like the internet. With the launch of ChatGPT-4 Turbo in November 2023, we developed a ChatGPT-based teaching application (https://chat.openai.com/g/g-1imx1py4K-chatge-medical-imaging) and integrated it into our undergraduat… ▽ More The integration of artificial intelligence (AI) chatbots into higher education marks a shift towards a new generation of pedagogical tools, mirroring the arrival of milestones like the internet. With the launch of ChatGPT-4 Turbo in November 2023, we developed a ChatGPT-based teaching application (https://chat.openai.com/g/g-1imx1py4K-chatge-medical-imaging) and integrated it into our undergraduate medical imaging course in the Spring 2024 semester. This study investigates the use of ChatGPT throughout a semester-long trial, providing insights into students' engagement, perception, and the overall educational effectiveness of the technology. We systematically collected and analyzed data concerning students' interaction with ChatGPT, focusing on their attitudes, concerns, and usage patterns. The findings indicate that ChatGPT offers significant advantages such as improved information access and increased interactivity, but its adoption is accompanied by concerns about the accuracy of the information provided and the necessity for well-defined guidelines to optimize its use. △ Less

Submitted 8 July, 2024; originally announced July 2024.

arXiv:2407.05681 [pdf]

Bulk high-temperature superconductivity in the high-pressure tetragonal phase of bilayer La2PrNi2O7

Authors: Ningning Wang, Gang Wang, Xiaoling Shen, Jun Hou, Jun Luo, Xiaoping Ma, Huaixin Yang, Lifen Shi, Jie Dou, Jie Feng, Jie Yang, Yunqing Shi, Zhian Ren, Hanming Ma, Pengtao Yang, Ziyi Liu, Yue Liu, Hua Zhang, Xiaoli Dong, Yuxin Wang, Kun Jiang, Jiangping Hu, Stuart Calder, Jiaqiang Yan, Jianping Sun , et al. (4 additional authors not shown)

Abstract: The Ruddlesden-Popper (R-P) bilayer nickelate, La3Ni2O7, was recently found to show signatures of high-temperature superconductivity (HTSC) at pressures above 14 GPa. Subsequent investigations achieved zero resistance in single- and poly-crystalline samples under hydrostatic pressure conditions. Yet, obvious diamagnetic signals, the other hallmark of superconductors, are still lacking owing to the… ▽ More The Ruddlesden-Popper (R-P) bilayer nickelate, La3Ni2O7, was recently found to show signatures of high-temperature superconductivity (HTSC) at pressures above 14 GPa. Subsequent investigations achieved zero resistance in single- and poly-crystalline samples under hydrostatic pressure conditions. Yet, obvious diamagnetic signals, the other hallmark of superconductors, are still lacking owing to the filamentary nature with low superconducting volume fraction. The presence of a novel "1313" polymorph and competing R-P phases obscured proper identification of the phase for HTSC. Thus, achieving bulk HTSC and identifying the phase at play are the most prominent tasks at present. Here, we address these issues in the praseodymium (Pr)-doped La2PrNi2O7 polycrystalline samples. We find that the substitutions of Pr for La effectively inhibits the intergrowth of different R-P phases, resulting in nearly pure bilayer structure. For La2PrNi2O7, pressure-induced orthorhombic-to-tetragonal structural transition takes place at Pc ~ 11 GPa, above which HTSC emerges gradually upon further compression. The superconducting transition temperatures at 18-20 GPa reach Tconset = 82.5 K and Tczero = 60 K, which are the highest values among known nickelate superconductors. More importantly, bulk HTSC was testified by detecting clear diamagnetic signals below ~75 K corresponding to an estimated superconducting volume fraction ~ 57(5)% at 20 GPa. Our results not only resolve the existing controversies but also illuminate directions for exploring bulk HTSC in the bilayer nickelates. △ Less

Submitted 8 July, 2024; originally announced July 2024.

arXiv:2407.05548 [pdf, other]

Ferromagnetic inter-layer coupling in FeSe$_{1-x}$S$_{x}$ superconductors revealed by inelastic neutron scattering

Authors: Mingwei Ma, Philippe Bourges, Yvan Sidis, Jinzhao Sun, Guoqing Wang, Kazuki Iida, Kazuya Kamazawa, Jitae T. Park, Frederic Bourdarot, Zhian Ren, Yuan Li

Abstract: FeSe$_{1-x}$S$_{x}$ superconductors are commonly considered layered van der Waals materials with negligible inter-layer coupling. Here, using inelastic neutron scattering to study spin excitations in single-crystal samples, we reveal that the magnetic coupling between adjacent Fe layers is not only significant, as it affects excitations up to \textcolor{black}{15} meV, but also ferromagnetic in na… ▽ More FeSe$_{1-x}$S$_{x}$ superconductors are commonly considered layered van der Waals materials with negligible inter-layer coupling. Here, using inelastic neutron scattering to study spin excitations in single-crystal samples, we reveal that the magnetic coupling between adjacent Fe layers is not only significant, as it affects excitations up to \textcolor{black}{15} meV, but also ferromagnetic in nature, making the system different from most unconventional superconductors including iron pnictides. Our observation provides a new standpoint to understand the absence of magnetic order in FeSe$_{1-x}$S$_{x}$. Since intercalating between the Fe layers is known to enhance superconductivity and suppress the inter-layer coupling, superconductivity appears to be a more robust phenomenon in the two-dimensional limit than antiferromagnetic order. △ Less

Submitted 7 July, 2024; originally announced July 2024.

arXiv:2407.05298 [pdf, other]

Probing axion-like particles in leptonic decays of heavy mesons

Authors: Gang Yang, Tianhong Wang, Guo-Li Wang

Abstract: We study the possibility to find the axion-like particles (ALPs) through the leptonic decays of heavy mesons. There are some deviations between the Standard Model (SM) predictions of the branching ratios of the leptonic decays of mesons and the experimental data. This provides some space for the existence of decay channels where the ALP is one of the products. Three scenarios are considered: first… ▽ More We study the possibility to find the axion-like particles (ALPs) through the leptonic decays of heavy mesons. There are some deviations between the Standard Model (SM) predictions of the branching ratios of the leptonic decays of mesons and the experimental data. This provides some space for the existence of decay channels where the ALP is one of the products. Three scenarios are considered: first, the ALP is only coupled to one single charged fermion, namely, the quark, the antiquark, or the charged lepton; second, the ALP is only coupled to quark and antiquark with the same strength; third, the ALP is coupled to all the charged fermions with the same strength. The constraints of the coupling strength in different scenarios are obtained by comparing the experimental data of the branching ratios of leptonic decays of $B^-$, $D^+$, and $D_s^+$ mesons with the theoretical predictions which are achieved by using the Bethe-Salpeter (BS) method. These constraints are further applied to predict the upper limits of the leptonic decay processes of the $B_c^-$ meson in which the ALP participates. △ Less

Submitted 7 July, 2024; originally announced July 2024.

Comments: 17 pages, 23 figures

arXiv:2407.04938 [pdf, other]

SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation

Authors: Guoan Wang, Jin Ye, Junlong Cheng, Tianbin Li, Zhaolin Chen, Jianfei Cai, Junjun He, Bohan Zhuang

Abstract: Volumetric medical image segmentation is pivotal in enhancing disease diagnosis, treatment planning, and advancing medical research. While existing volumetric foundation models for medical image segmentation, such as SAM-Med3D and SegVol, have shown remarkable performance on general organs and tumors, their ability to segment certain categories in clinical downstream tasks remains limited. Supervi… ▽ More Volumetric medical image segmentation is pivotal in enhancing disease diagnosis, treatment planning, and advancing medical research. While existing volumetric foundation models for medical image segmentation, such as SAM-Med3D and SegVol, have shown remarkable performance on general organs and tumors, their ability to segment certain categories in clinical downstream tasks remains limited. Supervised Finetuning (SFT) serves as an effective way to adapt such foundation models for task-specific downstream tasks but at the cost of degrading the general knowledge previously stored in the original foundation model.To address this, we propose SAM-Med3D-MoE, a novel framework that seamlessly integrates task-specific finetuned models with the foundational model, creating a unified model at minimal additional training expense for an extra gating network. This gating network, in conjunction with a selection strategy, allows the unified model to achieve comparable performance of the original models in their respective tasks both general and specialized without updating any parameters of them.Our comprehensive experiments demonstrate the efficacy of SAM-Med3D-MoE, with an average Dice performance increase from 53 to 56.4 on 15 specific classes. It especially gets remarkable gains of 29.6, 8.5, 11.2 on the spinal cord, esophagus, and right hip, respectively. Additionally, it achieves 48.9 Dice on the challenging SPPIN2023 Challenge, significantly surpassing the general expert's performance of 32.3. We anticipate that SAM-Med3D-MoE can serve as a new framework for adapting the foundation model to specific areas in medical image analysis. Codes and datasets will be publicly available. △ Less

Submitted 5 July, 2024; originally announced July 2024.

Journal ref: MICCAI 2024

arXiv:2407.04365 [pdf, ps, other]

Simulation of Spin Chains with off-diagonal Coupling Using Inchworm Method

Authors: Yixiao Sun, Geshuo Wang, Zhenning Cai

Abstract: We study the dynamical simulation of open quantum spin chain with nearest neighboring coupling, where each spin in the chain is associated with a harmonic bath. This is an extension of our previous work [G. Wang and Z. Cai, J. Chem. Theory Comput., 19, 8523--8540, 2023] by generalizing the application of the inchworm method and the technique of modular path integrals from diagonally coupled cases… ▽ More We study the dynamical simulation of open quantum spin chain with nearest neighboring coupling, where each spin in the chain is associated with a harmonic bath. This is an extension of our previous work [G. Wang and Z. Cai, J. Chem. Theory Comput., 19, 8523--8540, 2023] by generalizing the application of the inchworm method and the technique of modular path integrals from diagonally coupled cases to off-diagonally coupled cases. Additionally, to reduce computational and memory cost in long time simulation, we apply tensor-train representation to efficiently represent the reduced density matrix of the spin chains, and employ the transfer tensor method (TTM) to avoid exponential growth of computational cost with respect to time. Abundant numerical experiments are performed to validate our method. △ Less

Submitted 5 July, 2024; originally announced July 2024.

arXiv:2407.03222 [pdf, ps, other]

On $δ$-Stable Minimal Hypersurfaces in $\mathbb{R}^{n+1}$

Authors: Han Hong, Haizhong Li, Gaoming Wang

Abstract: In this paper, we extend several results established for stable minimal hypersurfaces to $δ$-stable minimal hypersurfaces. These include the regularity and compactness theorems for immersed $δ$-stable minimal hypersurfaces in $\mathbb{R}^{n+1}$ when $n \geq 3$ and $δ> \frac{n-2}{n}$, as well as the $δ$-stable Bernstein theorem for $n=3$ and $n=4$ for properly immersion. The range of $δ$ is optimal… ▽ More In this paper, we extend several results established for stable minimal hypersurfaces to $δ$-stable minimal hypersurfaces. These include the regularity and compactness theorems for immersed $δ$-stable minimal hypersurfaces in $\mathbb{R}^{n+1}$ when $n \geq 3$ and $δ> \frac{n-2}{n}$, as well as the $δ$-stable Bernstein theorem for $n=3$ and $n=4$ for properly immersion. The range of $δ$ is optimal, as the $n$-dimensional catenoid in $\mathbb{R}^{n+1}$ is $\frac{n-2}{n}$-stable. △ Less

Submitted 3 July, 2024; originally announced July 2024.

Comments: 37 pages. Comments are welcome

arXiv:2407.02893 [pdf, other]

An Uncertainty-guided Tiered Self-training Framework for Active Source-free Domain Adaptation in Prostate Segmentation

Authors: Zihao Luo, Xiangde Luo, Zijun Gao, Guotai Wang

Abstract: Deep learning models have exhibited remarkable efficacy in accurately delineating the prostate for diagnosis and treatment of prostate diseases, but challenges persist in achieving robust generalization across different medical centers. Source-free Domain Adaptation (SFDA) is a promising technique to adapt deep segmentation models to address privacy and security concerns while reducing domain shif… ▽ More Deep learning models have exhibited remarkable efficacy in accurately delineating the prostate for diagnosis and treatment of prostate diseases, but challenges persist in achieving robust generalization across different medical centers. Source-free Domain Adaptation (SFDA) is a promising technique to adapt deep segmentation models to address privacy and security concerns while reducing domain shifts between source and target domains. However, recent literature indicates that the performance of SFDA remains far from satisfactory due to unpredictable domain gaps. Annotating a few target domain samples is acceptable, as it can lead to significant performance improvement with a low annotation cost. Nevertheless, due to extremely limited annotation budgets, careful consideration is needed in selecting samples for annotation. Inspired by this, our goal is to develop Active Source-free Domain Adaptation (ASFDA) for medical image segmentation. Specifically, we propose a novel Uncertainty-guided Tiered Self-training (UGTST) framework, consisting of efficient active sample selection via entropy-based primary local peak filtering to aggregate global uncertainty and diversity-aware redundancy filter, coupled with a tiered self-learning strategy, achieves stable domain adaptation. Experimental results on cross-center prostate MRI segmentation datasets revealed that our method yielded marked advancements, with a mere 5% annotation, exhibiting an average Dice score enhancement of 9.78% and 7.58% in two target domains compared with state-of-the-art methods, on par with fully supervised learning. Code is available at:https://github.com/HiLab-git/UGTST △ Less

Submitted 4 July, 2024; v1 submitted 3 July, 2024; originally announced July 2024.

Comments: 11 pages, 3 figures, 2 tables, accept to MICCAI 2024

arXiv:2407.02639 [pdf, other]

Holistically-Nested Structure-Aware Graph Neural Network for Road Extraction

Authors: Tinghuai Wang, Guangming Wang, Kuan Eeik Tan

Abstract: Convolutional neural networks (CNN) have made significant advances in detecting roads from satellite images. However, existing CNN approaches are generally repurposed semantic segmentation architectures and suffer from the poor delineation of long and curved regions. Lack of overall road topology and structure information further deteriorates their performance on challenging remote sensing images.… ▽ More Convolutional neural networks (CNN) have made significant advances in detecting roads from satellite images. However, existing CNN approaches are generally repurposed semantic segmentation architectures and suffer from the poor delineation of long and curved regions. Lack of overall road topology and structure information further deteriorates their performance on challenging remote sensing images. This paper presents a novel multi-task graph neural network (GNN) which simultaneously detects both road regions and road borders; the inter-play between these two tasks unlocks superior performance from two perspectives: (1) the hierarchically detected road borders enable the network to capture and encode holistic road structure to enhance road connectivity (2) identifying the intrinsic correlation of semantic landcover regions mitigates the difficulty in recognizing roads cluttered by regions with similar appearance. Experiments on challenging dataset demonstrate that the proposed architecture can improve the road border delineation and road extraction accuracy compared with the existing methods. △ Less

Submitted 8 July, 2024; v1 submitted 2 July, 2024; originally announced July 2024.

arXiv:2407.02261 [pdf, other]

Federated Distillation for Medical Image Classification: Towards Trustworthy Computer-Aided Diagnosis

Authors: Sufen Ren, Yule Hu, Shengchao Chen, Guanjun Wang

Abstract: Medical image classification plays a crucial role in computer-aided clinical diagnosis. While deep learning techniques have significantly enhanced efficiency and reduced costs, the privacy-sensitive nature of medical imaging data complicates centralized storage and model training. Furthermore, low-resource healthcare organizations face challenges related to communication overhead and efficiency du… ▽ More Medical image classification plays a crucial role in computer-aided clinical diagnosis. While deep learning techniques have significantly enhanced efficiency and reduced costs, the privacy-sensitive nature of medical imaging data complicates centralized storage and model training. Furthermore, low-resource healthcare organizations face challenges related to communication overhead and efficiency due to increasing data and model scales. This paper proposes a novel privacy-preserving medical image classification framework based on federated learning to address these issues, named FedMIC. The framework enables healthcare organizations to learn from both global and local knowledge, enhancing local representation of private data despite statistical heterogeneity. It provides customized models for organizations with diverse data distributions while minimizing communication overhead and improving efficiency without compromising performance. Our FedMIC enhances robustness and practical applicability under resource-constrained conditions. We demonstrate FedMIC's effectiveness using four public medical image datasets for classical medical image classification tasks. △ Less

Submitted 3 July, 2024; v1 submitted 2 July, 2024; originally announced July 2024.

Comments: Work in progress. This paper is the first to introduce intra-client knowledge distillation in the context of trustworthy medical image classification. arXiv admin note: text overlap with arXiv:2401.01493

arXiv:2407.02165 [pdf, other]

WildAvatar: Web-scale In-the-wild Video Dataset for 3D Avatar Creation

Authors: Zihao Huang, Shoukang Hu, Guangcong Wang, Tianqi Liu, Yuhang Zang, Zhiguo Cao, Wei Li, Ziwei Liu

Abstract: Existing human datasets for avatar creation are typically limited to laboratory environments, wherein high-quality annotations (e.g., SMPL estimation from 3D scans or multi-view images) can be ideally provided. However, their annotating requirements are impractical for real-world images or videos, posing challenges toward real-world applications on current avatar creation methods. To this end, we… ▽ More Existing human datasets for avatar creation are typically limited to laboratory environments, wherein high-quality annotations (e.g., SMPL estimation from 3D scans or multi-view images) can be ideally provided. However, their annotating requirements are impractical for real-world images or videos, posing challenges toward real-world applications on current avatar creation methods. To this end, we propose the WildAvatar dataset, a web-scale in-the-wild human avatar creation dataset extracted from YouTube, with $10,000+$ different human subjects and scenes. WildAvatar is at least $10\times$ richer than previous datasets for 3D human avatar creation. We evaluate several state-of-the-art avatar creation methods on our dataset, highlighting the unexplored challenges in real-world applications on avatar creation. We also demonstrate the potential for generalizability of avatar creation methods, when provided with data at scale. We publicly release our data source links and annotations, to push forward 3D human avatar creation and other related fields for real-world applications. △ Less

Submitted 10 July, 2024; v1 submitted 2 July, 2024; originally announced July 2024.

Comments: Project page: https://wildavatar.github.io/

arXiv:2407.01527 [pdf, other]

KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches

Authors: Jiayi Yuan, Hongyi Liu, Shaochen, Zhong, Yu-Neng Chuang, Songchen Li, Guanchu Wang, Duy Le, Hongye Jin, Vipin Chaudhary, Zhaozhuo Xu, Zirui Liu, Xia Hu

Abstract: Long context capability is a crucial competency for large language models (LLMs) as it mitigates the human struggle to digest long-form texts. This capability enables complex task-solving scenarios such as book summarization, code assistance, and many more tasks that are traditionally manpower-intensive. However, transformer-based LLMs face significant challenges with long context input due to the… ▽ More Long context capability is a crucial competency for large language models (LLMs) as it mitigates the human struggle to digest long-form texts. This capability enables complex task-solving scenarios such as book summarization, code assistance, and many more tasks that are traditionally manpower-intensive. However, transformer-based LLMs face significant challenges with long context input due to the growing size of the KV cache and the intrinsic complexity of attending to extended inputs; where multiple schools of efficiency-driven approaches -- such as KV cache quantization, token dropping, prompt compression, linear-time sequence models, and hybrid architectures -- have been proposed to produce efficient yet long context-capable models. Despite these advancements, no existing work has comprehensively benchmarked these methods in a reasonably aligned environment. In this work, we fill this gap by providing a taxonomy of current methods and evaluating 10+ state-of-the-art approaches across seven categories of long context tasks. Our work reveals numerous previously unknown phenomena and offers insights -- as well as a friendly workbench -- for the future development of long context-capable LLMs. The source code will be available at https://github.com/henryzhongsc/longctx_bench △ Less

Submitted 1 July, 2024; originally announced July 2024.

arXiv:2407.00470 [pdf]

Unusual Pore Volume Dependence of Water Sorption in Monolithic Metal-Organic Framework

Authors: Jiawang Li, Guang Wang, Hongzhao Fan, Zhigang Li, Chi Yan Tso, Yanguang Zhou

Abstract: Monolithic metal-organic frameworks (MOFs), which have a continuous structure composed of small primary MOF particles and amorphous networks, are demonstrated to possess larger pore volume and thus better larger gas uptake capacity compared to their powder forms. Here, we systematically investigated the water vapor adsorption kinetics in a prototypical MOF, i.e., MOF-801. Our results show that the… ▽ More Monolithic metal-organic frameworks (MOFs), which have a continuous structure composed of small primary MOF particles and amorphous networks, are demonstrated to possess larger pore volume and thus better larger gas uptake capacity compared to their powder forms. Here, we systematically investigated the water vapor adsorption kinetics in a prototypical MOF, i.e., MOF-801. Our results show that the total pore volume (average pore diameter) of the monolithic MOF-801 is 0.831 cm3/g (5.20 nm) which is much larger than that of powder MOF-801, i.e., 0.488 cm3/g (1.95 nm). Unexpectedly, we find that the water uptake capacity of monolithic MOF-801 is much lower than that of powder MOF-801 when the RH ranges from 10% to 90%. Our molecular dynamics simulations further demonstrate that the unexpected water uptake capacity of monolithic MOF-801 at RH of 10%~90% is caused by the water film formed by the capillary condensation in these mesopores of monolithic MOF-801. The water molecules can overcome the capillary force when the RH is higher than 90%, and then leads to the increase of the corresponding water uptake capacity of monolithic MOF-801. Our findings reveal the underlying mechanisms for water adsorption kinetics in both powder and monolithic MOFs, which could motivate and benefit the new passive cooling or water harvesting system design based on MOFs. △ Less

Submitted 29 June, 2024; originally announced July 2024.

arXiv:2407.00296 [pdf]

SolarSAM: Building-scale Photovoltaic Potential Assessment Based on Segment Anything Model (SAM) and Remote Sensing for Emerging City

Authors: Guohao Wang

Abstract: Driven by advancements in photovoltaic (PV) technology, solar energy has emerged as a promising renewable energy source, due to its ease of integration onto building rooftops, facades, and windows. For the emerging cities, the lack of detailed street-level data presents a challenge for effectively assessing the potential of building-integrated photovoltaic (BIPV). To address this, this study intro… ▽ More Driven by advancements in photovoltaic (PV) technology, solar energy has emerged as a promising renewable energy source, due to its ease of integration onto building rooftops, facades, and windows. For the emerging cities, the lack of detailed street-level data presents a challenge for effectively assessing the potential of building-integrated photovoltaic (BIPV). To address this, this study introduces SolarSAM, a novel BIPV evaluation method that leverages remote sensing imagery and deep learning techniques, and an emerging city in northern China is utilized to validate the model performance. During the process, SolarSAM segmented various building rooftops using text prompt guided semantic segmentation. Separate PV models were then developed for Rooftop PV, Facade-integrated PV, and PV windows systems, using this segmented data and local climate information. The potential for BIPV installation, solar power generation, and city-wide power self-sufficiency were assessed, revealing that the annual BIPV power generation potential surpassed the city's total electricity consumption by a factor of 2.5. Economic and environmental analysis were also conducted, including levelized cost of electricity and carbon reduction calculations, comparing different BIPV systems across various building categories. These findings demonstrated the model's performance and reveled the potential of BIPV power generation in the future. △ Less

Submitted 28 June, 2024; originally announced July 2024.

arXiv:2407.00285 [pdf, other]

Imaging of single barium atoms in a second matrix site in solid xenon for barium tagging in a $^{136}$Xe double beta decay experiment

Authors: M. Yvaine, D. Fairbank, J. Soderstrom, C. Taylor, J. Stanley, T. Walton, C. Chambers, A. Iverson, W. Fairbank, S. Al Kharusi, A. Amy, E. Angelico, A. Anker, I. J. Arnquist, A. Atencio, J. Bane, V. Belov, E. P. Bernard, T. Bhatta, A. Bolotnikov, J. Breslin, P. A. Breur, J. P. Brodsky, E. Brown, T. Brunner , et al. (112 additional authors not shown)

Abstract: Neutrinoless double beta decay is one of the most sensitive probes for new physics beyond the Standard Model of particle physics. One of the isotopes under investigation is $^{136}$Xe, which would double beta decay into $^{136}$Ba. Detecting the single $^{136}$Ba daughter provides a sort of ultimate tool in the discrimination against backgrounds. Previous work demonstrated the ability to perform s… ▽ More Neutrinoless double beta decay is one of the most sensitive probes for new physics beyond the Standard Model of particle physics. One of the isotopes under investigation is $^{136}$Xe, which would double beta decay into $^{136}$Ba. Detecting the single $^{136}$Ba daughter provides a sort of ultimate tool in the discrimination against backgrounds. Previous work demonstrated the ability to perform single atom imaging of Ba atoms in a single-vacancy site of a solid xenon matrix. In this paper, the effort to identify signal from individual barium atoms is extended to Ba atoms in a hexa-vacancy site in the matrix and is achieved despite increased photobleaching in this site. Abrupt fluorescence turn-off of a single Ba atom is also observed. Significant recovery of fluorescence signal lost through photobleaching is demonstrated upon annealing of Ba deposits in the Xe ice. Following annealing, it is observed that Ba atoms in the hexa-vacancy site exhibit antibleaching while Ba atoms in the tetra-vacancy site exhibit bleaching. This may be evidence for a matrix site transfer upon laser excitation. Our findings offer a path of continued research toward tagging of Ba daughters in all significant sites in solid xenon. △ Less

Submitted 28 June, 2024; originally announced July 2024.

Comments: 9 pages, 8 figures

arXiv:2407.00046 [pdf, other]

Barrier-Augmented Lagrangian for GPU-based Elastodynamic Contact

Authors: Dewen Guo, Minchen Li, Yin Yang, Guoping Wang, Sheng Li

Abstract: We propose a GPU-based iterative method for accelerated elastodynamic simulation with the log-barrier-based contact model. While Newton's method is a conventional choice for solving the interior-point system, the presence of ill-conditioned log barriers often necessitates a direct solution at each linearized substep and costs substantial storage and computational overhead. Moreover, constraint set… ▽ More We propose a GPU-based iterative method for accelerated elastodynamic simulation with the log-barrier-based contact model. While Newton's method is a conventional choice for solving the interior-point system, the presence of ill-conditioned log barriers often necessitates a direct solution at each linearized substep and costs substantial storage and computational overhead. Moreover, constraint sets that vary in each iteration present additional challenges in algorithm convergence. Our method employs a novel barrier-augmented Lagrangian method to improve system conditioning and solver efficiency by adaptively updating an augmentation constraint sets. This enables the utilization of a scalable, inexact Newton-PCG solver with sparse GPU storage, eliminating the need for direct factorization. We further enhance PCG convergence speed with a domain-decomposed warm start strategy based on an eigenvalue spectrum approximated through our in-time assembly. Demonstrating significant scalability improvements, our method makes simulations previously impractical on 128 GB of CPU memory feasible with only 8 GB of GPU memory and orders-of-magnitude faster. Additionally, our method adeptly handles stiff problems, surpassing the capabilities of existing GPU-based interior-point methods. Our results, validated across various complex collision scenarios involving intricate geometries and large deformations, highlight the exceptional performance of our approach. △ Less

Submitted 4 June, 2024; originally announced July 2024.

Comments: 17 pages, 30 figures

arXiv:2406.19844 [pdf, other]

StreamMOTP: Streaming and Unified Framework for Joint 3D Multi-Object Tracking and Trajectory Prediction

Authors: Jiaheng Zhuang, Guoan Wang, Siyu Zhang, Xiyang Wang, Hangning Zhou, Ziyao Xu, Chi Zhang, Zhiheng Li

Abstract: 3D multi-object tracking and trajectory prediction are two crucial modules in autonomous driving systems. Generally, the two tasks are handled separately in traditional paradigms and a few methods have started to explore modeling these two tasks in a joint manner recently. However, these approaches suffer from the limitations of single-frame training and inconsistent coordinate representations bet… ▽ More 3D multi-object tracking and trajectory prediction are two crucial modules in autonomous driving systems. Generally, the two tasks are handled separately in traditional paradigms and a few methods have started to explore modeling these two tasks in a joint manner recently. However, these approaches suffer from the limitations of single-frame training and inconsistent coordinate representations between tracking and prediction tasks. In this paper, we propose a streaming and unified framework for joint 3D Multi-Object Tracking and trajectory Prediction (StreamMOTP) to address the above challenges. Firstly, we construct the model in a streaming manner and exploit a memory bank to preserve and leverage the long-term latent features for tracked objects more effectively. Secondly, a relative spatio-temporal positional encoding strategy is introduced to bridge the gap of coordinate representations between the two tasks and maintain the pose-invariance for trajectory prediction. Thirdly, we further improve the quality and consistency of predicted trajectories with a dual-stream predictor. We conduct extensive experiments on popular nuSences dataset and the experimental results demonstrate the effectiveness and superiority of StreamMOTP, which outperforms previous methods significantly on both tasks. Furthermore, we also prove that the proposed framework has great potential and advantages in actual applications of autonomous driving. △ Less

Submitted 28 June, 2024; originally announced June 2024.

arXiv:2406.19646 [pdf, other]

Time-optimal Flight in Cluttered Environments via Safe Reinforcement Learning

Authors: Wei Xiao, Zhaohan Feng, Ziyu Zhou, Jian Sun, Gang Wang, Jie Chen

Abstract: This paper addresses the problem of guiding a quadrotor through a predefined sequence of waypoints in cluttered environments, aiming to minimize the flight time while avoiding collisions. Previous approaches either suffer from prolonged computational time caused by solving complex non-convex optimization problems or are limited by the inherent smoothness of polynomial trajectory representations, t… ▽ More This paper addresses the problem of guiding a quadrotor through a predefined sequence of waypoints in cluttered environments, aiming to minimize the flight time while avoiding collisions. Previous approaches either suffer from prolonged computational time caused by solving complex non-convex optimization problems or are limited by the inherent smoothness of polynomial trajectory representations, thereby restricting the flexibility of movement. In this work, we present a safe reinforcement learning approach for autonomous drone racing with time-optimal flight in cluttered environments. The reinforcement learning policy, trained using safety and terminal rewards specifically designed to enforce near time-optimal and collision-free flight, outperforms current state-of-the-art algorithms. Additionally, experimental results demonstrate the efficacy of the proposed approach in achieving both minimum flight time and obstacle avoidance objectives in complex environments, with a commendable $66.7\%$ success rate in unseen, challenging settings. △ Less

Submitted 28 June, 2024; originally announced June 2024.

Comments: 7 pages, 3 figures,

arXiv:2406.19623 [pdf]

FRA-DiagSys: A Transformer Winding Fault Diagnosis System for Identifying Fault Types and degrees Using Frequency Response Analysis

Authors: Guohao Wang

Abstract: The electric power transformer is a critical component in electrical distribution networks, and the diagnosis of faults in transformers is an important research area. Frequency Response Analysis (FRA) methods are widely used for analyzing winding faults in transformers, particularly in Chinese power stations. However, the current approach relies on manual expertise to interpret FRA curves, which c… ▽ More The electric power transformer is a critical component in electrical distribution networks, and the diagnosis of faults in transformers is an important research area. Frequency Response Analysis (FRA) methods are widely used for analyzing winding faults in transformers, particularly in Chinese power stations. However, the current approach relies on manual expertise to interpret FRA curves, which can be both skill-intensive and lacks precision. This study presents a novel approach using a Multilayer perceptron model to directly model and analyze FRA data, simulating various winding fault types and degrees in 12-disc winding and 10-disc winding transformers with different connection configurations, resulting in three distinct datasets. Six different Multilayer perceptron architectures were developed, with optimal models achieving recognition accuracies of over 99.7% for diagnosing fault degrees and more than 90% for fault types. Hence, this paper has yielded a model architecture that exhibits commendable performance in diagnosing various fault types and their severities in different models of transformers when utilizing different FRA connection methods. Additionally, a specialized diagnostic system called FRA-DiagSys with two-stage model utilization was developed, achieving 100% accuracy in diagnosing fault types and degrees for a specific winding-10 power transformer, surpassing other diagnostic methods and strategies. △ Less

Submitted 27 June, 2024; originally announced June 2024.

arXiv:2406.18884 [pdf, other]

Sequential three-way group decision-making for double hierarchy hesitant fuzzy linguistic term set

Authors: Nanfang Luo, Qinghua Zhang, Qin Xie, Yutai Wang, Longjun Yin, Guoyin Wang

Abstract: Group decision-making (GDM) characterized by complexity and uncertainty is an essential part of various life scenarios. Most existing researches lack tools to fuse information quickly and interpret decision results for partially formed decisions. This limitation is particularly noticeable when there is a need to improve the efficiency of GDM. To address this issue, a novel multi-level sequential t… ▽ More Group decision-making (GDM) characterized by complexity and uncertainty is an essential part of various life scenarios. Most existing researches lack tools to fuse information quickly and interpret decision results for partially formed decisions. This limitation is particularly noticeable when there is a need to improve the efficiency of GDM. To address this issue, a novel multi-level sequential three-way decision for group decision-making (S3W-GDM) method is constructed from the perspective of granular computing. This method simultaneously considers the vagueness, hesitation, and variation of GDM problems under double hierarchy hesitant fuzzy linguistic term sets (DHHFLTS) environment. First, for fusing information efficiently, a novel multi-level expert information fusion method is proposed, and the concepts of expert decision table and the extraction/aggregation of decision-leveled information based on the multi-level granularity are defined. Second, the neighborhood theory, outranking relation and regret theory (RT) are utilized to redesign the calculations of conditional probability and relative loss function. Then, the granular structure of DHHFLTS based on the sequential three-way decision (S3WD) is defined to improve the decision-making efficiency, and the decision-making strategy and interpretation of each decision-level are proposed. Furthermore, the algorithm of S3W-GDM is given. Finally, an illustrative example of diagnosis is presented, and the comparative and sensitivity analysis with other methods are performed to verify the efficiency and rationality of the proposed method. △ Less

Submitted 27 June, 2024; originally announced June 2024.

arXiv:2406.18485 [pdf, other]

LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism

Authors: Diandian Gu, Peng Sun, Qinghao Hu, Ting Huang, Xun Chen, Yingtong Xiong, Guoteng Wang, Qiaoling Chen, Shangchun Zhao, Jiarui Fang, Yonggang Wen, Tianwei Zhang, Xin Jin, Xuanzhe Liu

Abstract: Efficiently training LLMs with long sequences is important yet challenged by the massive computation and memory requirements. Sequence parallelism has been proposed to tackle these problems, but existing methods suffer from scalability or efficiency issues. We propose LoongTrain, a novel system to efficiently train LLMs with long sequences at scale. The core of LoongTrain is the 2D-Attention mecha… ▽ More Efficiently training LLMs with long sequences is important yet challenged by the massive computation and memory requirements. Sequence parallelism has been proposed to tackle these problems, but existing methods suffer from scalability or efficiency issues. We propose LoongTrain, a novel system to efficiently train LLMs with long sequences at scale. The core of LoongTrain is the 2D-Attention mechanism, which combines both head-parallel and context-parallel techniques to break the scalability constraints while maintaining efficiency. We introduce Double-Ring-Attention and analyze the performance of device placement strategies to further speed up training. We implement LoongTrain with the hybrid ZeRO and Selective Checkpoint++ techniques. Experiment results show that LoongTrain outperforms state-of-the-art baselines, i.e., DeepSpeed-Ulysses and Megatron Context Parallelism, in both end-to-end training speed and scalability, and improves Model FLOPs Utilization (MFU) by up to 2.88x. △ Less

Submitted 26 June, 2024; originally announced June 2024.

arXiv:2406.18063 [pdf, other]

Data-driven imaging geometric recovery of ultrahigh resolution robotic micro-CT for in-vivo and other applications

Authors: Mengzhou Li, Guibin Zan, Wenbin Yun, Josef Uher, John Wen, Ge Wang

Abstract: We introduce an ultrahigh-resolution (50μm\) robotic micro-CT design for localized imaging of carotid plaques using robotic arms, cutting-edge detector, and machine learning technologies. To combat geometric error-induced artifacts in interior CT scans, we propose a data-driven geometry estimation method that maximizes the consistency between projection data and the reprojection counterparts of a… ▽ More We introduce an ultrahigh-resolution (50μm\) robotic micro-CT design for localized imaging of carotid plaques using robotic arms, cutting-edge detector, and machine learning technologies. To combat geometric error-induced artifacts in interior CT scans, we propose a data-driven geometry estimation method that maximizes the consistency between projection data and the reprojection counterparts of a reconstructed volume. Particularly, we use a normalized cross correlation metric to overcome the projection truncation effect. Our approach is validated on a robotic CT scan of a sacrificed mouse and a micro-CT phantom scan, both producing sharper images with finer details than that prior correction. △ Less

Submitted 26 June, 2024; originally announced June 2024.

Comments: 4-page paper for 8th International Conference on Computational and Mathematical Biomedical Engineering

arXiv:2406.17971 [pdf, other]

Robust integration of external control data in randomized trials

Authors: Rickard Karlsson, Guanbo Wang, Jesse H. Krijthe, Issa J. Dahabreh

Abstract: One approach for increasing the efficiency of randomized trials is the use of "external controls" -- individuals who received the control treatment in the trial during routine practice or in prior experimental studies. Existing external control methods, however, can have substantial bias if the populations underlying the trial and the external control data are not exchangeable. Here, we characteri… ▽ More One approach for increasing the efficiency of randomized trials is the use of "external controls" -- individuals who received the control treatment in the trial during routine practice or in prior experimental studies. Existing external control methods, however, can have substantial bias if the populations underlying the trial and the external control data are not exchangeable. Here, we characterize a randomization-aware class of treatment effect estimators in the population underlying the trial that remain consistent and asymptotically normal when using external control data, even when exchangeability does not hold. We consider two members of this class of estimators: the well-known augmented inverse probability weighting trial-only estimator, which is the efficient estimator when only trial data are used; and a more efficient member of the class when exchangeability holds and external control data are available, which we refer to as the optimized randomization-aware estimator. To achieve robust integration of external control data in trial analyses, we then propose a combined estimator based on the efficient trial-only estimator and the optimized randomization-aware estimator. We show that the combined estimator is consistent and no less efficient than the most efficient of the two component estimators, whether the exchangeability assumption holds or not. We examine the estimators' performance in simulations and we illustrate their use with data from two trials of paliperidone extended-release for schizophrenia. △ Less

Submitted 25 June, 2024; originally announced June 2024.

arXiv:2406.17685 [pdf, other]

CMBFSCNN: Cosmic Microwave Background Polarization Foreground Subtraction with Convolutional Neural Network

Authors: Ye-Peng Yan, Si-Yu Li, Guo-Jian Wang, Zirui Zhang, Jun-Qing Xia

Abstract: In our previous study, we introduced a machine-learning technique, namely CMBFSCNN, for the removal of foreground contamination in cosmic microwave background (CMB) polarization data. This method was successfully employed on actual observational data from the Planck mission. In this study, we extend our investigation by considering the CMB lensing effect in simulated data and utilizing the CMBFSCN… ▽ More In our previous study, we introduced a machine-learning technique, namely CMBFSCNN, for the removal of foreground contamination in cosmic microwave background (CMB) polarization data. This method was successfully employed on actual observational data from the Planck mission. In this study, we extend our investigation by considering the CMB lensing effect in simulated data and utilizing the CMBFSCNN approach to recover the CMB lensing B-mode power spectrum from multi-frequency observational maps. Our method is first applied to simulated data with the performance of CMB-S4 experiment. We achieve reliable recovery of the noisy CMB Q (or U) maps with a mean absolute difference of $0.016\pm0.008\ μ$K (or $0.021\pm0.002\ μ$K) for the CMB-S4 experiment. To address the residual instrumental noise in the foreground-cleaned map, we employ a "half-split maps" approach, where the entire dataset is divided into two segments sharing the same sky signal but having uncorrelated noise. Using cross-correlation techniques between two recovered half-split maps, we effectively reduce instrumental noise effects at the power spectrum level. As a result, we achieve precise recovery of the CMB EE and lensing B-mode power spectra. Furthermore, we also extend our pipeline to full-sky simulated data with the performance of LiteBIRD experiment. As expected, various foregrounds are cleanly removed from the foreground contamination observational maps, and recovered EE and lensing B-mode power spectra exhibit excellent agreement with the true results. Finally, we discuss the dependency of our method on the foreground models. △ Less

Submitted 25 June, 2024; originally announced June 2024.

Comments: 26 pages, 16 figures, 3 table, accepted by ApJS

arXiv:2406.17507 [pdf, other]

ACE: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling

Authors: Minghui Fang, Shengpeng Ji, Jialong Zuo, Hai Huang, Yan Xia, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu, Gang Wang, Zhenhua Dong, Zhou Zhao

Abstract: Generative retrieval, which has demonstrated effectiveness in text-to-text retrieval, utilizes a sequence-to-sequence model to directly generate candidate identifiers based on natural language queries. Without explicitly computing the similarity between queries and candidates, generative retrieval surpasses dual-tower models in both speed and accuracy on large-scale corpora, providing new insights… ▽ More Generative retrieval, which has demonstrated effectiveness in text-to-text retrieval, utilizes a sequence-to-sequence model to directly generate candidate identifiers based on natural language queries. Without explicitly computing the similarity between queries and candidates, generative retrieval surpasses dual-tower models in both speed and accuracy on large-scale corpora, providing new insights for cross-modal retrieval. However, constructing identifiers for multimodal data remains an untapped problem, and the modality gap between natural language queries and multimodal candidates hinders retrieval performance due to the absence of additional encoders. To this end, we propose a pioneering generAtive Cross-modal rEtrieval framework (ACE), which is a comprehensive framework for end-to-end cross-modal retrieval based on coarse-to-fine semantic modeling. We propose combining K-Means and RQ-VAE to construct coarse and fine tokens, serving as identifiers for multimodal data. Correspondingly, we design the coarse-to-fine feature fusion strategy to efficiently align natural language queries and candidate identifiers. ACE is the first work to comprehensively demonstrate the feasibility of generative approach on text-to-image/audio/video retrieval, challenging the dominance of the embedding-based dual-tower architecture. Extensive experiments show that ACE achieves state-of-the-art performance in cross-modal retrieval and outperforms the strong baselines on Recall@1 by 15.27% on average. △ Less

Submitted 25 June, 2024; originally announced June 2024.

arXiv:2406.17054 [pdf, other]

Mean-Field Langevin Dynamics for Signed Measures via a Bilevel Approach

Authors: Guillaume Wang, Alireza Mousavi-Hosseini, Lénaïc Chizat

Abstract: Mean-field Langevin dynamics (MLFD) is a class of interacting particle methods that tackle convex optimization over probability measures on a manifold, which are scalable, versatile, and enjoy computational guarantees. However, some important problems -- such as risk minimization for infinite width two-layer neural networks, or sparse deconvolution -- are originally defined over the set of signed,… ▽ More Mean-field Langevin dynamics (MLFD) is a class of interacting particle methods that tackle convex optimization over probability measures on a manifold, which are scalable, versatile, and enjoy computational guarantees. However, some important problems -- such as risk minimization for infinite width two-layer neural networks, or sparse deconvolution -- are originally defined over the set of signed, rather than probability, measures. In this paper, we investigate how to extend the MFLD framework to convex optimization problems over signed measures. Among two known reductions from signed to probability measures -- the lifting and the bilevel approaches -- we show that the bilevel reduction leads to stronger guarantees and faster rates (at the price of a higher per-iteration complexity). In particular, we investigate the convergence rate of MFLD applied to the bilevel reduction in the low-noise regime and obtain two results. First, this dynamics is amenable to an annealing schedule, adapted from Suzuki et al. (2023), that results in improved convergence rates to a fixed multiplicative accuracy. Second, we investigate the problem of learning a single neuron with the bilevel approach and obtain local exponential convergence rates that depend polynomially on the dimension and noise level (to compare with the exponential dependence that would result from prior analyses). △ Less

Submitted 26 June, 2024; v1 submitted 24 June, 2024; originally announced June 2024.

Comments: 57 pages, 1 figure

arXiv:2406.17006 [pdf, other]

Probing the nature of the $χ_{c1}(3872)$ state using radiative decays

Authors: LHCb collaboration, R. Aaij, A. S. W. Abdelmotteleb, C. Abellan Beteta, F. Abudinén, T. Ackernley, A. A. Adefisoye, B. Adeva, M. Adinolfi, P. Adlarson, C. Agapopoulou, C. A. Aidala, Z. Ajaltouni, S. Akar, K. Akiba, P. Albicocco, J. Albrecht, F. Alessio, M. Alexander, Z. Aliouche, P. Alvarez Cartelle, R. Amalric, S. Amato, J. L. Amey, Y. Amhis , et al. (1094 additional authors not shown)

Abstract: The radiative decays $χ_{c1}(3872)\rightarrowψ(2S)γ$ and $χ_{c1}(3872)\rightarrow J/ψγ$ are used to probe the~nature of the~$χ_{c1}(3872)$ state using proton-proton collision data collected with the LHCb detector, corresponding to an~integrated luminosity of~9fb$^{-1}$. Using the~$B^+\rightarrow χ_{c1}(3872)K^+$decay, the $χ_{c1}(3872)\rightarrow ψ(2S)γ$ process is observed for the first time and… ▽ More The radiative decays $χ_{c1}(3872)\rightarrowψ(2S)γ$ and $χ_{c1}(3872)\rightarrow J/ψγ$ are used to probe the~nature of the~$χ_{c1}(3872)$ state using proton-proton collision data collected with the LHCb detector, corresponding to an~integrated luminosity of~9fb$^{-1}$. Using the~$B^+\rightarrow χ_{c1}(3872)K^+$decay, the $χ_{c1}(3872)\rightarrow ψ(2S)γ$ process is observed for the first time and the ratio of its partial width to that of the $χ_{c1}(3872)\rightarrow J/ψγ$ decay is measured to be $$ \frac{Γ_{χ_{c1}(3872)\rightarrow ψ(2S)γ}} {Γ_{χ_{c1}(3872)\rightarrow J/ψγ}} = 1.67 \pm 0.21 \pm 0.12 \pm0.04 , $$ where the first uncertainty is statistical, the second systematic and the third is due to the uncertainties on the branching fractions of the $ψ(2S)$ and $J/ψ$ mesons. The measured ratio makes the interpretation of the $χ_{c1}(3872)$ state as a~pure $D^0\bar{D}^{*0}+\bar{D}^0D^{*0}$ molecule questionable and strongly indicates a sizeable compact charmonium or tetraquark component within the $χ_{c1}(3872)$ state. △ Less

Submitted 24 June, 2024; originally announced June 2024.

Comments: 31 pages, 2 figures. All figures and tables, along with any supplementary material and additional information, are available at https://cern.ch/lhcbproject/Publications/p/LHCb-PAPER-2024-015.html (LHCb public pages)

Report number: LHCb-PAPER-2024-015, CERN-EP-2025-157

arXiv:2406.16989 [pdf, other]

Retrieval-Augmented Mixture of LoRA Experts for Uploadable Machine Learning

Authors: Ziyu Zhao, Leilei Gan, Guoyin Wang, Yuwei Hu, Tao Shen, Hongxia Yang, Kun Kuang, Fei Wu

Abstract: Low-Rank Adaptation (LoRA) offers an efficient way to fine-tune large language models (LLMs). Its modular and plug-and-play nature allows the integration of various domain-specific LoRAs, enhancing LLM capabilities. Open-source platforms like Huggingface and Modelscope have introduced a new computational paradigm, Uploadable Machine Learning (UML). In UML, contributors use decentralized data to tr… ▽ More Low-Rank Adaptation (LoRA) offers an efficient way to fine-tune large language models (LLMs). Its modular and plug-and-play nature allows the integration of various domain-specific LoRAs, enhancing LLM capabilities. Open-source platforms like Huggingface and Modelscope have introduced a new computational paradigm, Uploadable Machine Learning (UML). In UML, contributors use decentralized data to train specialized adapters, which are then uploaded to a central platform to improve LLMs. This platform uses these domain-specific adapters to handle mixed-task requests requiring personalized service. Previous research on LoRA composition either focuses on specific tasks or fixes the LoRA selection during training. However, in UML, the pool of LoRAs is dynamically updated with new uploads, requiring a generalizable selection mechanism for unseen LoRAs. Additionally, the mixed-task nature of downstream requests necessitates personalized services. To address these challenges, we propose Retrieval-Augmented Mixture of LoRA Experts (RAMoLE), a framework that adaptively retrieves and composes multiple LoRAs based on input prompts. RAMoLE has three main components: LoraRetriever for identifying and retrieving relevant LoRAs, an on-the-fly MoLE mechanism for coordinating the retrieved LoRAs, and efficient batch inference for handling heterogeneous requests. Experimental results show that RAMoLE consistently outperforms baselines, highlighting its effectiveness and scalability. △ Less

Submitted 24 June, 2024; originally announced June 2024.

Comments: arXiv admin note: substantial text overlap with arXiv:2402.09997

arXiv:2406.16637 [pdf, other]

A 100 Mpc$^2$ structure traced by hyperluminous galaxies around a massive $z$ = 2.85 protocluster

Authors: George C. P. Wang, Scott C. Chapman, Nikolaus Sulzenauer, Frank Bertoldi, Christopher C. Hayward, Ryley Hill, Satoshi Kikuta, Yuichi Matsuda, Douglas Rennehan, Douglas Scott, Ian Smail, Charles C. Steidel

Abstract: We present wide-field mapping at 850 $μ$m and 450 $μ$m of the $z$ = 2.85 protocluster in the HS1549$+$19 field using the Submillimetre Common User Bolometer Array 2 (SCUBA-2). Spectroscopic follow-up of 18 bright sources selected at 850 $μ$m, using the Nothern Extended Millimeter Array (NOEMA) and Atacama Large Millimeter Array (ALMA), confirms the majority lies near $z$ $\sim$ 2.85 and are likely… ▽ More We present wide-field mapping at 850 $μ$m and 450 $μ$m of the $z$ = 2.85 protocluster in the HS1549$+$19 field using the Submillimetre Common User Bolometer Array 2 (SCUBA-2). Spectroscopic follow-up of 18 bright sources selected at 850 $μ$m, using the Nothern Extended Millimeter Array (NOEMA) and Atacama Large Millimeter Array (ALMA), confirms the majority lies near $z$ $\sim$ 2.85 and are likely members of the structure. Interpreting the spectroscopic redshifts as distance measurements, we find that the SMGs span 90 Mpc$^2$ in the plane of the sky and demarcate a 4100 Mpc$^3$ "pancake"-shaped structure in three dimensions. We find that the high star-formation rates (SFRs) of these SMGs result in a total SFR of 20,000 M$_\odot$ yr$^{-1}$ only from the brightest galaxies in the protocluster. These rapidly star-forming SMGs can be interpreted as massive galaxies growing rapidly at large cluster-centric distances before collapsing into a virialized structure. We find that the SMGs trace the Lyman-$α$ surface density profile. Comparison with simulations suggests that HS1549$+$19 could be building a structure comparable to the most massive clusters in the present-day Universe. △ Less

Submitted 24 June, 2024; originally announced June 2024.

arXiv:2406.16268 [pdf, other]

Efficient Antagonistic k-plex Enumeration in Signed Graphs

Authors: Lantian Xu, Rong-Hua Li, Dong Wen, Qiangqiang Dai, Guoren Wang, Lu Qin

Abstract: A signed graph is a graph where each edge receives a sign, positive or negative. The signed graph model has been used in many real applications, such as protein complex discovery and social network analysis. Finding cohesive subgraphs in signed graphs is a fundamental problem. A k-plex is a common model for cohesive subgraphs in which every vertex is adjacent to all but at most k vertices within t… ▽ More A signed graph is a graph where each edge receives a sign, positive or negative. The signed graph model has been used in many real applications, such as protein complex discovery and social network analysis. Finding cohesive subgraphs in signed graphs is a fundamental problem. A k-plex is a common model for cohesive subgraphs in which every vertex is adjacent to all but at most k vertices within the subgraph. In this paper, we propose the model of size-constrained antagonistic k-plex in a signed graph. The proposed model guarantees that the resulting subgraph is a k-plex and can be divided into two sub-k-plexes, both of which have positive inner edges and negative outer edges. This paper aims to identify all maximal antagonistic k-plexes in a signed graph. Through rigorous analysis, we show that the problem is NP-Hardness. We propose a novel framework for maximal antagonistic k-plexes utilizing set enumeration. Efficiency is improved through pivot pruning and early termination based on the color bound. Preprocessing techniques based on degree and dichromatic graphs effectively narrow the search space before enumeration. Extensive experiments on real-world datasets demonstrate our algorithm's efficiency, effectiveness, and scalability. △ Less

Submitted 23 June, 2024; originally announced June 2024.

arXiv:2406.15768 [pdf, other]

MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception

Authors: Guanqun Wang, Xinyu Wei, Jiaming Liu, Ray Zhang, Yichi Zhang, Kevin Zhang, Maurice Chong, Shanghang Zhang

Abstract: In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant strides in perception tasks, such as detection and segmentation. However, MLLMs mainly focus on high-level image-text interpretations and struggle with fine-grained visual understanding,… ▽ More In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant strides in perception tasks, such as detection and segmentation. However, MLLMs mainly focus on high-level image-text interpretations and struggle with fine-grained visual understanding, and vision perception models usually suffer from open-world distribution shifts due to their limited model capacity. To overcome these challenges, we propose the Mutually Reinforced Multimodal Large Language Model (MR-MLLM), a novel framework that synergistically enhances visual perception and multimodal comprehension. First, a shared query fusion mechanism is proposed to harmonize detailed visual inputs from vision models with the linguistic depth of language models, enhancing multimodal comprehension and vision perception synergistically. Second, we propose the perception-enhanced cross-modal integration method, incorporating novel modalities from vision perception outputs, like object detection bounding boxes, to capture subtle visual elements, thus enriching the understanding of both visual and textual data. In addition, an innovative perception-embedded prompt generation mechanism is proposed to embed perceptual information into the language model's prompts, aligning the responses contextually and perceptually for a more accurate multimodal interpretation. Extensive experiments demonstrate MR-MLLM's superior performance in various multimodal comprehension and vision perception tasks, particularly those requiring corner case vision perception and fine-grained language comprehension. △ Less

Submitted 22 June, 2024; originally announced June 2024.

Comments: 14 pages, 8 figures

arXiv:2406.15222 [pdf]

Rapid and Accurate Diagnosis of Acute Aortic Syndrome using Non-contrast CT: A Large-scale, Retrospective, Multi-center and AI-based Study

Authors: Yujian Hu, Yilang Xiang, Yan-Jie Zhou, Yangyan He, Shifeng Yang, Xiaolong Du, Chunlan Den, Youyao Xu, Gaofeng Wang, Zhengyao Ding, Jingyong Huang, Wenjun Zhao, Xuejun Wu, Donglin Li, Qianqian Zhu, Zhenjiang Li, Chenyang Qiu, Ziheng Wu, Yunjun He, Chen Tian, Yihui Qiu, Zuodong Lin, Xiaolong Zhang, Yuan He, Zhenpeng Yuan , et al. (15 additional authors not shown)

Abstract: Chest pain symptoms are highly prevalent in emergency departments (EDs), where acute aortic syndrome (AAS) is a catastrophic cardiovascular emergency with a high fatality rate, especially when timely and accurate treatment is not administered. However, current triage practices in the ED can cause up to approximately half of patients with AAS to have an initially missed diagnosis or be misdiagnosed… ▽ More Chest pain symptoms are highly prevalent in emergency departments (EDs), where acute aortic syndrome (AAS) is a catastrophic cardiovascular emergency with a high fatality rate, especially when timely and accurate treatment is not administered. However, current triage practices in the ED can cause up to approximately half of patients with AAS to have an initially missed diagnosis or be misdiagnosed as having other acute chest pain conditions. Subsequently, these AAS patients will undergo clinically inaccurate or suboptimal differential diagnosis. Fortunately, even under these suboptimal protocols, nearly all these patients underwent non-contrast CT covering the aorta anatomy at the early stage of differential diagnosis. In this study, we developed an artificial intelligence model (DeepAAS) using non-contrast CT, which is highly accurate for identifying AAS and provides interpretable results to assist in clinical decision-making. Performance was assessed in two major phases: a multi-center retrospective study (n = 20,750) and an exploration in real-world emergency scenarios (n = 137,525). In the multi-center cohort, DeepAAS achieved a mean area under the receiver operating characteristic curve of 0.958 (95% CI 0.950-0.967). In the real-world cohort, DeepAAS detected 109 AAS patients with misguided initial suspicion, achieving 92.6% (95% CI 76.2%-97.5%) in mean sensitivity and 99.2% (95% CI 99.1%-99.3%) in mean specificity. Our AI model performed well on non-contrast CT at all applicable early stages of differential diagnosis workflows, effectively reduced the overall missed diagnosis and misdiagnosis rate from 48.8% to 4.8% and shortened the diagnosis time for patients with misguided initial suspicion from an average of 681.8 (74-11,820) mins to 68.5 (23-195) mins. DeepAAS could effectively fill the gap in the current clinical workflow without requiring additional tests. △ Less

Submitted 24 June, 2024; v1 submitted 13 June, 2024; originally announced June 2024.

Comments: under peer review

arXiv:2406.14173 [pdf, ps, other]

Time-Delay Interferometry for ASTROD-GW

Authors: Gang Wang

Abstract: In the detection of gravitational waves in space, the arm lengths between spacecraft are not equal due to their orbital motion. Consequently, the equal arm length Michelson interferometer used in Earth laboratories is not suitable for space. To achieve the necessary sensitivity for space gravitational wave detectors, laser frequency noise must be suppressed below secondary noise sources such as op… ▽ More In the detection of gravitational waves in space, the arm lengths between spacecraft are not equal due to their orbital motion. Consequently, the equal arm length Michelson interferometer used in Earth laboratories is not suitable for space. To achieve the necessary sensitivity for space gravitational wave detectors, laser frequency noise must be suppressed below secondary noise sources such as optical path noise and acceleration noise. To suppress laser frequency noise, time-delay interferometry (TDI) is employed to match the two optical paths and retain gravitational wave signals. Since planets and other solar system bodies perturb the orbits of spacecraft and affect TDI performance, we simulate the time delay numerically using the CGC2.7 ephemeris framework. To examine the feasibility of TDI for the ASTROD-GW mission, we devised a set of 10-year and a set of 20-year optimized mission orbits for the three spacecraft starting on June 21, 2028, and calculated the path mismatches in the first- and second-generation TDI channels. The results demonstrate that all second-generation TDI channels meet the ASTROD-GW requirements. A geometric approach is used in the analysis and synthesis of both first-generation and second-generation TDI to clearly illustrate the construction process. △ Less

Submitted 20 June, 2024; originally announced June 2024.

Comments: This master's thesis was originally written in Chinese and submitted in 2011. This version is quickly translated with the assistance of ChatGPT. In Chapter 7, a second-generation TDI bank was developed, and some of the TDI observables are related to recent works (arXiv:2403.01490 and arXiv:2406.11305). Comments and feedback are welcome

arXiv:2406.14088 [pdf, other]

ReaLHF: Optimized RLHF Training for Large Language Models through Parameter Reallocation

Authors: Zhiyu Mei, Wei Fu, Kaiwei Li, Guangju Wang, Huanchen Zhang, Yi Wu

Abstract: Reinforcement Learning from Human Feedback (RLHF) stands as a pivotal technique in empowering large language model (LLM) applications. Since RLHF involves diverse computational workloads and intricate dependencies among multiple LLMs, directly adopting parallelization techniques from supervised training can result in sub-optimal performance. To overcome this limitation, we propose a novel approach… ▽ More Reinforcement Learning from Human Feedback (RLHF) stands as a pivotal technique in empowering large language model (LLM) applications. Since RLHF involves diverse computational workloads and intricate dependencies among multiple LLMs, directly adopting parallelization techniques from supervised training can result in sub-optimal performance. To overcome this limitation, we propose a novel approach named parameter ReaLlocation, which dynamically redistributes LLM parameters in the cluster and adapts parallelization strategies during training. Building upon this idea, we introduce ReaLHF, a pioneering system capable of automatically discovering and running efficient execution plans for RLHF training given the desired algorithmic and hardware configurations. ReaLHF formulates the execution plan for RLHF as an augmented dataflow graph. Based on this formulation, ReaLHF employs a tailored search algorithm with a lightweight cost estimator to discover an efficient execution plan. Subsequently, the runtime engine deploys the selected plan by effectively parallelizing computations and redistributing parameters. We evaluate ReaLHF on the LLaMA-2 models with up to $4\times70$ billion parameters and 128 GPUs. The experiment results showcase ReaLHF's substantial speedups of $2.0-10.6\times$ compared to baselines. Furthermore, the execution plans generated by ReaLHF exhibit an average of $26\%$ performance improvement over heuristic approaches based on Megatron-LM. The source code of ReaLHF is publicly available at https://github.com/openpsi-project/ReaLHF . △ Less

Submitted 20 June, 2024; originally announced June 2024.

Comments: 13 pages (15 pages with references), 13 figures

arXiv:2406.14045 [pdf, other]

Understanding Different Design Choices in Training Large Time Series Models

Authors: Yu-Neng Chuang, Songchen Li, Jiayi Yuan, Guanchu Wang, Kwei-Herng Lai, Leisheng Yu, Sirui Ding, Chia-Yuan Chang, Qiaoyu Tan, Daochen Zha, Xia Hu

Abstract: Inspired by Large Language Models (LLMs), Time Series Forecasting (TSF), a long-standing task in time series analysis, is undergoing a transition towards Large Time Series Models (LTSMs), aiming to train universal transformer-based models for TSF. However, training LTSMs on heterogeneous time series data poses unique challenges, including diverse frequencies, dimensions, and patterns across datase… ▽ More Inspired by Large Language Models (LLMs), Time Series Forecasting (TSF), a long-standing task in time series analysis, is undergoing a transition towards Large Time Series Models (LTSMs), aiming to train universal transformer-based models for TSF. However, training LTSMs on heterogeneous time series data poses unique challenges, including diverse frequencies, dimensions, and patterns across datasets. Recent endeavors have studied and evaluated various design choices aimed at enhancing LTSM training and generalization capabilities, spanning pre-processing techniques, model configurations, and dataset configurations. In this work, we comprehensively analyze these design choices and aim to identify the best practices for training LTSM. Moreover, we propose \emph{time series prompt}, a novel statistical prompting strategy tailored to time series data. Furthermore, based on the observations in our analysis, we introduce \texttt{LTSM-bundle}, which bundles the best design choices we have identified. Empirical results demonstrate that \texttt{LTSM-bundle} achieves superior zero-shot and few-shot performances compared to state-of-the-art LSTMs and traditional TSF methods on benchmark datasets. △ Less

Submitted 20 June, 2024; originally announced June 2024.

arXiv:2406.13998 [pdf, other]

Transversal Hamilton paths and cycles

Authors: Yangyang Cheng, Wanting Sun, Guanghui Wang, Lan Wei

Abstract: Given a collection $\mathcal{G} =\{G_1,G_2,\dots,G_m\}$ of graphs on the common vertex set $V$ of size $n$, an $m$-edge graph $H$ on the same vertex set $V$ is transversal in $\mathcal{G}$ if there exists a bijection $\varphi :E(H)\rightarrow [m]$ such that $e \in E(G_{\varphi(e)})$ for all $e\in E(H)$. Denote $δ(\mathcal{G}):=\operatorname*{min}\left\{δ(G_i): i\in [m]\right\}$. In this paper, we… ▽ More Given a collection $\mathcal{G} =\{G_1,G_2,\dots,G_m\}$ of graphs on the common vertex set $V$ of size $n$, an $m$-edge graph $H$ on the same vertex set $V$ is transversal in $\mathcal{G}$ if there exists a bijection $\varphi :E(H)\rightarrow [m]$ such that $e \in E(G_{\varphi(e)})$ for all $e\in E(H)$. Denote $δ(\mathcal{G}):=\operatorname*{min}\left\{δ(G_i): i\in [m]\right\}$. In this paper, we first establish a minimum degree condition for the existence of transversal Hamilton paths in $\mathcal{G}$: if $n=m+1$ and $δ(\mathcal{G})\geq \frac{n-1}{2}$, then $\mathcal{G}$ contains a transversal Hamilton path. This solves a problem proposed by [Li, Li and Li, J. Graph Theory, 2023]. As a continuation of the transversal version of Dirac's theorem [Joos and Kim, Bull. Lond. Math. Soc., 2020] and the stability result for transversal Hamilton cycles [Cheng and Staden, arXiv:2403.09913v1], our second result characterizes all graph collections with minimum degree at least $\frac{n}{2}-1$ and without transversal Hamilton cycles. We obtain an analogous result for transversal Hamilton paths. The proof is a combination of the stability result for transversal Hamilton paths or cycles, transversal blow-up lemma, along with some structural analysis. △ Less

Submitted 20 June, 2024; originally announced June 2024.

Comments: 33 pages, 10 figures

MSC Class: 05C35

arXiv:2406.13768 [pdf, other]

FastPersist: Accelerating Model Checkpointing in Deep Learning

Authors: Guanhua Wang, Olatunji Ruwase, Bing Xie, Yuxiong He

Abstract: Model checkpoints are critical Deep Learning (DL) artifacts that enable fault tolerance for training and downstream applications, such as inference. However, writing checkpoints to persistent storage, and other I/O aspects of DL training, are mostly ignored by compute-focused optimization efforts for faster training of rapidly growing models and datasets. Towards addressing this imbalance, we prop… ▽ More Model checkpoints are critical Deep Learning (DL) artifacts that enable fault tolerance for training and downstream applications, such as inference. However, writing checkpoints to persistent storage, and other I/O aspects of DL training, are mostly ignored by compute-focused optimization efforts for faster training of rapidly growing models and datasets. Towards addressing this imbalance, we propose FastPersist to accelerate checkpoint creation in DL training. FastPersist combines three novel techniques: (i) NVMe optimizations for faster checkpoint writes to SSDs, (ii) efficient write parallelism using the available SSDs in training environments, and (iii) overlapping checkpointing with independent training computations. Our evaluation using real world dense and sparse DL models shows that FastPersist creates checkpoints in persistent storage up to 116x faster than baseline, and enables per-iteration checkpointing with negligible overhead. △ Less

Submitted 19 June, 2024; originally announced June 2024.

Comments: 11 pages

arXiv:2406.13674 [pdf, other]

Rethinking Abdominal Organ Segmentation (RAOS) in the clinical scenario: A robustness evaluation benchmark with challenging cases

Authors: Xiangde Luo, Zihan Li, Shaoting Zhang, Wenjun Liao, Guotai Wang

Abstract: Deep learning has enabled great strides in abdominal multi-organ segmentation, even surpassing junior oncologists on common cases or organs. However, robustness on corner cases and complex organs remains a challenging open problem for clinical adoption. To investigate model robustness, we collected and annotated the RAOS dataset comprising 413 CT scans ($\sim$80k 2D images, $\sim$8k 3D organ annot… ▽ More Deep learning has enabled great strides in abdominal multi-organ segmentation, even surpassing junior oncologists on common cases or organs. However, robustness on corner cases and complex organs remains a challenging open problem for clinical adoption. To investigate model robustness, we collected and annotated the RAOS dataset comprising 413 CT scans ($\sim$80k 2D images, $\sim$8k 3D organ annotations) from 413 patients each with 17 (female) or 19 (male) labelled organs, manually delineated by oncologists. We grouped scans based on clinical information into 1) diagnosis/radiotherapy (317 volumes), 2) partial excision without the whole organ missing (22 volumes), and 3) excision with the whole organ missing (74 volumes). RAOS provides a potential benchmark for evaluating model robustness including organ hallucination. It also includes some organs that can be very hard to access on public datasets like the rectum, colon, intestine, prostate and seminal vesicles. We benchmarked several state-of-the-art methods in these three clinical groups to evaluate performance and robustness. We also assessed cross-generalization between RAOS and three public datasets. This dataset and comprehensive analysis establish a potential baseline for future robustness research: \url{https://github.com/Luoxd1996/RAOS}. △ Less

Submitted 19 June, 2024; originally announced June 2024.

Comments: 10 pages, 1 figure, 6 tables, Early Accept to MICCAI 2024

arXiv:2406.13672 [pdf, other]

Q-SNNs: Quantized Spiking Neural Networks

Authors: Wenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao, Honglin Cao, Zhenbang Ren, Guoqing Wang, Malu Zhang, Yang Yang

Abstract: Brain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to represent information and process them in an asynchronous event-driven manner, offering an energy-efficient paradigm for the next generation of machine intelligence. However, the current focus within the SNN community prioritizes accuracy optimization through the development of large-scale models, limiting their viability in r… ▽ More Brain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to represent information and process them in an asynchronous event-driven manner, offering an energy-efficient paradigm for the next generation of machine intelligence. However, the current focus within the SNN community prioritizes accuracy optimization through the development of large-scale models, limiting their viability in resource-constrained and low-power edge devices. To address this challenge, we introduce a lightweight and hardware-friendly Quantized SNN (Q-SNN) that applies quantization to both synaptic weights and membrane potentials. By significantly compressing these two key elements, the proposed Q-SNNs substantially reduce both memory usage and computational complexity. Moreover, to prevent the performance degradation caused by this compression, we present a new Weight-Spike Dual Regulation (WS-DR) method inspired by information entropy theory. Experimental evaluations on various datasets, including static and neuromorphic, demonstrate that our Q-SNNs outperform existing methods in terms of both model size and accuracy. These state-of-the-art results in efficiency and efficacy suggest that the proposed method can significantly improve edge intelligent computing. △ Less

Submitted 19 June, 2024; originally announced June 2024.

Comments: 8 pages, 5 figures

arXiv:2406.13157 [pdf, other]

Genetics-based deperturbation analysis for the spin-orbit coupled ${\rm A}^1Σ^+$ and ${\rm b}^3Π_{0^+}$ states of LiRb

Authors: Yide Yin, Xuhui Bai, Xuechun Li, Xin-Yu Luo, Jie Yu, Gaoren Wang, Yongchang Han

Abstract: We present a deperturbation analysis of the spin-orbit coupled $\rm A^1Σ^+$ and $\rm b^3Π_{0^+}$ states of LiRb based on the rovibrational energy levels observed previously by photoassociation spectroscopy in bosonic $^7$Li$^{85}$Rb molecule. Using the genetic algorithm, we fit the potential energy curves of the $\rm A^1Σ^+$ state and the $\rm b^3Π$ state into point-wise form. We then fit these po… ▽ More We present a deperturbation analysis of the spin-orbit coupled $\rm A^1Σ^+$ and $\rm b^3Π_{0^+}$ states of LiRb based on the rovibrational energy levels observed previously by photoassociation spectroscopy in bosonic $^7$Li$^{85}$Rb molecule. Using the genetic algorithm, we fit the potential energy curves of the $\rm A^1Σ^+$ state and the $\rm b^3Π$ state into point-wise form. We then fit these point-wise potentials along with the spin-orbit coupling into expanded Morse oscillator functional form and optimise analytical parameters based on the experimental data. From the fitted results, we calculate the transition dipole moment matrix elements for transitions from the rovibrational levels of the coupled $\rm A^1Σ^+$-$\rm b^3Π_{0^+}$ state to the Feshbach state and the absolute rovibrational ground state for fermionic $^6$Li$^{87}$Rb molecule. Based on the calculated transition dipole moment matrix elements, several levels of the coupled $\rm A^1Σ^+$-$\rm b^3Π_{0^+}$ state are predicted to be suitable as the intermediate state for stimulated Raman adiabatic passage transfer from the Feshbach state to the absolute rovibrational ground state. In addition, we also provide a similar estimation for ${\rm B}^1Π$-${\rm c}^3Σ_1^+$-${\rm b}^3Π_1$ state based on available $ab\ initio$ interaction potentials. △ Less

Submitted 4 July, 2024; v1 submitted 18 June, 2024; originally announced June 2024.

Comments: 12 pages, 9 figures

arXiv:2406.12111 [pdf, other]

Precision measurement of the $Ξ^-_b$ baryon lifetime

Authors: LHCb collaboration, R. Aaij, A. S. W. Abdelmotteleb, C. Abellan Beteta, F. Abudinén, T. Ackernley, A. A. Adefisoye, B. Adeva, M. Adinolfi, P. Adlarson, C. Agapopoulou, C. A. Aidala, Z. Ajaltouni, S. Akar, K. Akiba, P. Albicocco, J. Albrecht, F. Alessio, M. Alexander, Z. Aliouche, P. Alvarez Cartelle, R. Amalric, S. Amato, J. L. Amey, Y. Amhis , et al. (1064 additional authors not shown)

Abstract: A sample of $pp$ collision data, corresponding to an integrated luminosity of 5.5 fb$^{-1}$ and collected by the LHCb experiment during Run 2, is used to measure the ratio of the lifetime of the $Ξ^-_b$ baryon to that of the $Λ^0_b$ baryon, $r_τ\equivτ_{Ξ^-_b}/τ_{Λ^0_b}$. The value ${r_τ^{\rm Run\,2}=1.076\pm0.013\pm0.006}$ is obtained, where the first uncertainty is statistical and the second sys… ▽ More A sample of $pp$ collision data, corresponding to an integrated luminosity of 5.5 fb$^{-1}$ and collected by the LHCb experiment during Run 2, is used to measure the ratio of the lifetime of the $Ξ^-_b$ baryon to that of the $Λ^0_b$ baryon, $r_τ\equivτ_{Ξ^-_b}/τ_{Λ^0_b}$. The value ${r_τ^{\rm Run\,2}=1.076\pm0.013\pm0.006}$ is obtained, where the first uncertainty is statistical and the second systematic. This value is averaged with the corresponding value from Run 1 to obtain ${r_τ^{\rm Run\,1,2} = 1.078\pm0.012\pm0.007}$. Multiplying by the world-average value of the $Λ^0_b$ lifetime yields $τ_{Ξ^-_b}^{\rm Run~1,2} = 1.578\pm0.018\pm0.010\pm0.011$ ps, where the uncertainties are statistical, systematic, and due to the limited knowledge of the $Λ^0_b$ lifetime. This measurement improves the precision of the current world average of the $Ξ^-_b$ lifetime by about a factor of two, and is in good agreement with the most recent theoretical predictions. △ Less

Submitted 17 June, 2024; originally announced June 2024.

Comments: 12 pages, 5 figures. All figures and tables, along with any supplementary material and additional information, are available at https://cern.ch/lhcbproject/Publications/p/LHCb-PAPER-2014-010.html (LHCb public pages)

Report number: LHCb-PAPER-2024-010, CERN-EP-2024-139

arXiv:2406.11799 [pdf, other]

Mix-Domain Contrastive Learning for Unpaired H&E-to-IHC Stain Translation

Authors: Song Wang, Zhong Zhang, Huan Yan, Ming Xu, Guanghui Wang

Abstract: H&E-to-IHC stain translation techniques offer a promising solution for precise cancer diagnosis, especially in low-resource regions where there is a shortage of health professionals and limited access to expensive equipment. Considering the pixel-level misalignment of H&E-IHC image pairs, current research explores the pathological consistency between patches from the same positions of the image pa… ▽ More H&E-to-IHC stain translation techniques offer a promising solution for precise cancer diagnosis, especially in low-resource regions where there is a shortage of health professionals and limited access to expensive equipment. Considering the pixel-level misalignment of H&E-IHC image pairs, current research explores the pathological consistency between patches from the same positions of the image pair. However, most of them overemphasize the correspondence between domains or patches, overlooking the side information provided by the non-corresponding objects. In this paper, we propose a Mix-Domain Contrastive Learning (MDCL) method to leverage the supervision information in unpaired H&E-to-IHC stain translation. Specifically, the proposed MDCL method aggregates the inter-domain and intra-domain pathology information by estimating the correlation between the anchor patch and all the patches from the matching images, encouraging the network to learn additional contrastive knowledge from mixed domains. With the mix-domain pathology information aggregation, MDCL enhances the pathological consistency between the corresponding patches and the component discrepancy of the patches from the different positions of the generated IHC image. Extensive experiments on two H&E-to-IHC stain translation datasets, namely MIST and BCI, demonstrate that the proposed method achieves state-of-the-art performance across multiple metrics. △ Less

Submitted 17 June, 2024; originally announced June 2024.

arXiv:2406.11489 [pdf]

Transverse optical gradient force in untethered rotating metaspinners

Authors: Einstom Engay, Mahdi Shanei, Vasilii Mylnikov, Gan Wang, Peter Johansson, Giovanni Volpe, Mikael Käll

Abstract: We introduce optical metasurfaces as components of ultracompact untethered microscopic metaspinners capable of efficient light-induced rotation in a liquid environment. Illuminated by weakly focused light, a metaspinner generates torque via photon recoil through the metasurfaces' ability to bend light towards high angles despite their sub-wavelength thickness, thereby creating orbital angular mome… ▽ More We introduce optical metasurfaces as components of ultracompact untethered microscopic metaspinners capable of efficient light-induced rotation in a liquid environment. Illuminated by weakly focused light, a metaspinner generates torque via photon recoil through the metasurfaces' ability to bend light towards high angles despite their sub-wavelength thickness, thereby creating orbital angular momentum. We find that a metaspinner is subject to an anomalous transverse optical gradient force that acts in concert with the classical gradient force. Consequently, when two or more metaspinners are trapped together in a laser beam, they collectively orbit the optical axis in the opposite direction to their spinning motion, in stark contrast to rotors coupled through hydrodynamic or mechanical interactions. The metaspinners delineated herein not only serve to illustrate the vast possibilities of utilizing optical metasurfaces for fundamental exploration of optical torques, but they also represent potential building-blocks of artificial active matter systems, light-driven micromachinery, and general-purpose optomechanical devices. △ Less

Submitted 17 June, 2024; originally announced June 2024.

Showing 1–50 of 4,213 results for author: Wang, G