subscribe to arXiv mailings

DG-PIC: Domain Generalized Point-In-Context Learning for Point Cloud Understanding

Authors: Jincen Jiang, Qianyu Zhou, Yuhang Li, Xuequan Lu, Meili Wang, Lizhuang Ma, Jian Chang, Jian Jun Zhang

Abstract: Recent point cloud understanding research suffers from performance drops on unseen data, due to the distribution shifts across different domains. While recent studies use Domain Generalization (DG) techniques to mitigate this by learning domain-invariant features, most are designed for a single task and neglect the potential of testing data. Despite In-Context Learning (ICL) showcasing multi-task… ▽ More Recent point cloud understanding research suffers from performance drops on unseen data, due to the distribution shifts across different domains. While recent studies use Domain Generalization (DG) techniques to mitigate this by learning domain-invariant features, most are designed for a single task and neglect the potential of testing data. Despite In-Context Learning (ICL) showcasing multi-task learning capability, it usually relies on high-quality context-rich data and considers a single dataset, and has rarely been studied in point cloud understanding. In this paper, we introduce a novel, practical, multi-domain multi-task setting, handling multiple domains and multiple tasks within one unified model for domain generalized point cloud understanding. To this end, we propose Domain Generalized Point-In-Context Learning (DG-PIC) that boosts the generalizability across various tasks and domains at testing time. In particular, we develop dual-level source prototype estimation that considers both global-level shape contextual and local-level geometrical structures for representing source domains and a dual-level test-time feature shifting mechanism that leverages both macro-level domain semantic information and micro-level patch positional relationships to pull the target data closer to the source ones during the testing. Our DG-PIC does not require any model updates during the testing and can handle unseen domains and multiple tasks, \textit{i.e.,} point cloud reconstruction, denoising, and registration, within one unified model. We also introduce a benchmark for this new setting. Comprehensive experiments demonstrate that DG-PIC outperforms state-of-the-art techniques significantly. △ Less

Submitted 11 July, 2024; originally announced July 2024.

Comments: Accepted to ECCV 2024

arXiv:2407.08366 [pdf, other]

An Economic Framework for 6-DoF Grasp Detection

Authors: Xiao-Ming Wu, Jia-Feng Cai, Jian-Jian Jiang, Dian Zheng, Yi-Lin Wei, Wei-Shi Zheng

Abstract: Robotic grasping in clutters is a fundamental task in robotic manipulation. In this work, we propose an economic framework for 6-DoF grasp detection, aiming to economize the resource cost in training and meanwhile maintain effective grasp performance. To begin with, we discover that the dense supervision is the bottleneck of current SOTA methods that severely encumbers the entire training overload… ▽ More Robotic grasping in clutters is a fundamental task in robotic manipulation. In this work, we propose an economic framework for 6-DoF grasp detection, aiming to economize the resource cost in training and meanwhile maintain effective grasp performance. To begin with, we discover that the dense supervision is the bottleneck of current SOTA methods that severely encumbers the entire training overload, meanwhile making the training difficult to converge. To solve the above problem, we first propose an economic supervision paradigm for efficient and effective grasping. This paradigm includes a well-designed supervision selection strategy, selecting key labels basically without ambiguity, and an economic pipeline to enable the training after selection. Furthermore, benefit from the economic supervision, we can focus on a specific grasp, and thus we devise a focal representation module, which comprises an interactive grasp head and a composite score estimation to generate the specific grasp more accurately. Combining all together, the EconomicGrasp framework is proposed. Our extensive experiments show that EconomicGrasp surpasses the SOTA grasp method by about 3AP on average, and with extremely low resource cost, for about 1/4 training time cost, 1/8 memory cost and 1/30 storage cost. Our code is available at https://github.com/iSEE-Laboratory/EconomicGrasp. △ Less

Submitted 11 July, 2024; originally announced July 2024.

Comments: 19 pages, 7 figures. Accepted in ECCV 2024!

arXiv:2407.07651 [pdf, other]

Study of the decay and production properties of $D_{s1}(2536)$ and $D_{s2}^*(2573)$

Authors: M. Ablikim, M. N. Achasov, P. Adlarson, O. Afedulidis, X. C. Ai, R. Aliberti, A. Amoroso, Q. An, Y. Bai, O. Bakina, I. Balossino, Y. Ban, H. -R. Bao, V. Batozskaya, K. Begzsuren, N. Berger, M. Berlowski, M. Bertani, D. Bettoni, F. Bianchi, E. Bianco, A. Bortone, I. Boyko, R. A. Briere, A. Brueggemann , et al. (645 additional authors not shown)

Abstract: The $e^+e^-\rightarrow D_s^+D_{s1}(2536)^-$ and $e^+e^-\rightarrow D_s^+D^*_{s2}(2573)^-$ processes are studied using data samples collected with the BESIII detector at center-of-mass energies from 4.530 to 4.946~GeV. The absolute branching fractions of $D_{s1}(2536)^- \rightarrow \bar{D}^{*0}K^-$ and $D_{s2}^*(2573)^- \rightarrow \bar{D}^0K^-$ are measured for the first time to be… ▽ More The $e^+e^-\rightarrow D_s^+D_{s1}(2536)^-$ and $e^+e^-\rightarrow D_s^+D^*_{s2}(2573)^-$ processes are studied using data samples collected with the BESIII detector at center-of-mass energies from 4.530 to 4.946~GeV. The absolute branching fractions of $D_{s1}(2536)^- \rightarrow \bar{D}^{*0}K^-$ and $D_{s2}^*(2573)^- \rightarrow \bar{D}^0K^-$ are measured for the first time to be $(35.9\pm 4.8\pm 3.5)\%$ and $(37.4\pm 3.1\pm 4.6)\%$, respectively. The measurements are in tension with predictions based on the assumption that the $D_{s1}(2536)$ and $D_{s2}^*(2573)$ are dominated by a bare $c\bar{s}$ component. The $e^+e^-\rightarrow D_s^+D_{s1}(2536)^-$ and $e^+e^-\rightarrow D_s^+D^*_{s2}(2573)^-$ cross sections are measured, and a resonant structure at around 4.6~GeV with a width of 50~MeV is observed for the first time with a statistical significance of $15σ$ in the $e^+e^-\rightarrow D_s^+D^*_{s2}(2573)^-$ process. It could be the $Y(4626)$ found by the Belle collaboration in the $D_s^+D_{s1}(2536)^{-}$ final state, since they have similar masses and widths. There is also evidence for a structure at around 4.75~GeV in both processes. △ Less

Submitted 10 July, 2024; originally announced July 2024.

arXiv:2407.07593 [pdf]

Observation of non-Abelian band topology without time-reversal symmetry

Authors: Yuze Hu, Mingyu Tong, Tian Jiang, Jian-hua Jiang, Hongsheng Chen, Yihao Yang

Abstract: Going beyond the conventional theory, non-Abelian band topology uncovers the global quantum geometry of Bloch bands with multiple gaps and thus unveil a new paradigm for topological physics. However, to date, all non-Abelian topological materials are restricted to systems with time-reversal symmetry (T). Here, starting from a Kagome lattice inspired by Haldane model and designer gyromagnetic photo… ▽ More Going beyond the conventional theory, non-Abelian band topology uncovers the global quantum geometry of Bloch bands with multiple gaps and thus unveil a new paradigm for topological physics. However, to date, all non-Abelian topological materials are restricted to systems with time-reversal symmetry (T). Here, starting from a Kagome lattice inspired by Haldane model and designer gyromagnetic photonic crystals (PhCs), we show that T breaking can lead to rich non-Abelian topological physics, particularly the emergence of multigap antichiral edge states. Simply changing the magnetic flux of the Kagome lattice, or in-situ tuning the local magnetic field of the gyromagnetic PhCs, can lead to the unconventional creation, braiding, merging, and splitting of non-Abelian charged band nodes, alongside with the direct manipulation of the multigap antichiral edge states. Particularly, the quadratic point can be split into four Dirac points, a phenomenon unique in T-broken systems. Our theoretical and experimental findings will inspire a new direction in the study of non-Abelian physics in T-broken systems and open an unprecedent pathway for topological manipulation of electromagnetic waves. △ Less

Submitted 10 July, 2024; originally announced July 2024.

arXiv:2407.06931 [pdf, other]

A Unified Approach to Multi-task Legged Navigation: Temporal Logic Meets Reinforcement Learning

Authors: Jesse Jiang, Samuel Coogan, Ye Zhao

Abstract: This study examines the problem of hopping robot navigation planning to achieve simultaneous goal-directed and environment exploration tasks. We consider a scenario in which the robot has mandatory goal-directed tasks defined using Linear Temporal Logic (LTL) specifications as well as optional exploration tasks represented using a reward function. Additionally, there exists uncertainty in the robo… ▽ More This study examines the problem of hopping robot navigation planning to achieve simultaneous goal-directed and environment exploration tasks. We consider a scenario in which the robot has mandatory goal-directed tasks defined using Linear Temporal Logic (LTL) specifications as well as optional exploration tasks represented using a reward function. Additionally, there exists uncertainty in the robot dynamics which results in motion perturbation. We first propose an abstraction of 3D hopping robot dynamics which enables high-level planning and a neural-network-based optimization for low-level control. We then introduce a Multi-task Product IMDP (MT-PIMDP) model of the system and tasks. We propose a unified control policy synthesis algorithm which enables both task-directed goal-reaching behaviors as well as task-agnostic exploration to learn perturbations and reward. We provide a formal proof of the trade-off induced by prioritizing either LTL or RL actions. We demonstrate our methods with simulation case studies in a 2D world navigation environment. △ Less

Submitted 9 July, 2024; originally announced July 2024.

Comments: 8 pages, 4 figures

arXiv:2407.06204 [pdf, other]

A Survey on Mixture of Experts

Authors: Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim, Jiayi Huang

Abstract: Large language models (LLMs) have garnered unprecedented advancements across diverse fields, ranging from natural language processing to computer vision and beyond. The prowess of LLMs is underpinned by their substantial model size, extensive and diverse datasets, and the vast computational power harnessed during training, all of which contribute to the emergent abilities of LLMs (e.g., in-context… ▽ More Large language models (LLMs) have garnered unprecedented advancements across diverse fields, ranging from natural language processing to computer vision and beyond. The prowess of LLMs is underpinned by their substantial model size, extensive and diverse datasets, and the vast computational power harnessed during training, all of which contribute to the emergent abilities of LLMs (e.g., in-context learning) that are not present in small models. Within this context, the mixture of experts (MoE) has emerged as an effective method for substantially scaling up model capacity with minimal computation overhead, gaining significant attention from academia and industry. Despite its growing prevalence, there lacks a systematic and comprehensive review of the literature on MoE. This survey seeks to bridge that gap, serving as an essential resource for researchers delving into the intricacies of MoE. We first briefly introduce the structure of the MoE layer, followed by proposing a new taxonomy of MoE. Next, we overview the core designs for various MoE models including both algorithmic and systemic aspects, alongside collections of available open-source implementations, hyperparameter configurations and empirical evaluations. Furthermore, we delineate the multifaceted applications of MoE in practice, and outline some potential directions for future research. To facilitate ongoing updates and the sharing of cutting-edge developments in MoE research, we have established a resource repository accessible at https://github.com/withinmiaov/A-Survey-on-Mixture-of-Experts. △ Less

Submitted 26 June, 2024; originally announced July 2024.

arXiv:2407.06196 [pdf, other]

Poetry2Image: An Iterative Correction Framework for Images Generated from Chinese Classical Poetry

Authors: Jing Jiang, Yiran Ling, Binzhu Li, Pengxiang Li, Junming Piao, Yu Zhang

Abstract: Text-to-image generation models often struggle with key element loss or semantic confusion in tasks involving Chinese classical poetry.Addressing this issue through fine-tuning models needs considerable training costs. Additionally, manual prompts for re-diffusion adjustments need professional knowledge. To solve this problem, we propose Poetry2Image, an iterative correction framework for images g… ▽ More Text-to-image generation models often struggle with key element loss or semantic confusion in tasks involving Chinese classical poetry.Addressing this issue through fine-tuning models needs considerable training costs. Additionally, manual prompts for re-diffusion adjustments need professional knowledge. To solve this problem, we propose Poetry2Image, an iterative correction framework for images generated from Chinese classical poetry. Utilizing an external poetry dataset, Poetry2Image establishes an automated feedback and correction loop, which enhances the alignment between poetry and image through image generation models and subsequent re-diffusion modifications suggested by large language models (LLM). Using a test set of 200 sentences of Chinese classical poetry, the proposed method--when integrated with five popular image generation models--achieves an average element completeness of 70.63%, representing an improvement of 25.56% over direct image generation. In tests of semantic correctness, our method attains an average semantic consistency of 80.09%. The study not only promotes the dissemination of ancient poetry culture but also offers a reference for similar non-fine-tuning methods to enhance LLM generation. △ Less

Submitted 15 June, 2024; originally announced July 2024.

Comments: 13 pages, 7 figures

arXiv:2407.05712 [pdf, other]

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

Authors: Jianwen Jiang, Gaojie Lin, Zhengkun Rong, Chao Liang, Yongming Zhu, Jiaqi Yang, Tianyun Zhong

Abstract: Existing neural head avatars methods have achieved significant progress in the image quality and motion range of portrait animation. However, these methods neglect the computational overhead, and to the best of our knowledge, none is designed to run on mobile devices. This paper presents MobilePortrait, a lightweight one-shot neural head avatars method that reduces learning complexity by integrati… ▽ More Existing neural head avatars methods have achieved significant progress in the image quality and motion range of portrait animation. However, these methods neglect the computational overhead, and to the best of our knowledge, none is designed to run on mobile devices. This paper presents MobilePortrait, a lightweight one-shot neural head avatars method that reduces learning complexity by integrating external knowledge into both the motion modeling and image synthesis, enabling real-time inference on mobile devices. Specifically, we introduce a mixed representation of explicit and implicit keypoints for precise motion modeling and precomputed visual features for enhanced foreground and background synthesis. With these two key designs and using simple U-Nets as backbones, our method achieves state-of-the-art performance with less than one-tenth the computational demand. It has been validated to reach speeds of over 100 FPS on mobile devices and support both video and audio-driven inputs. △ Less

Submitted 8 July, 2024; originally announced July 2024.

arXiv:2407.05673 [pdf, other]

doi 10.1093/pasj/psae043

A deep analysis for New Horizons' KBO search images

Authors: Fumi Yoshida, Toshifumi Yanagisawa, Takashi Ito, Hirohisa Kurosaki, Makoto Yoshikawa, Kohki Kamiya, Ji-an Jiang, Alan Stern, Wesley C. Fraser, Susan D. Benecchi, Anne J. Verbiscer

Abstract: Observation datasets acquired by the Hyper Suprime-Cam (HSC) on the Subaru Telescope for NASA's New Horizons mission target search were analyzed through a method devised by JAXA. The method makes use of Field Programmable Gate arrays and was originally used to detect fast-moving objects such as space debris or near-Earth asteroids. Here we present an application of the method to detect slow-moving… ▽ More Observation datasets acquired by the Hyper Suprime-Cam (HSC) on the Subaru Telescope for NASA's New Horizons mission target search were analyzed through a method devised by JAXA. The method makes use of Field Programmable Gate arrays and was originally used to detect fast-moving objects such as space debris or near-Earth asteroids. Here we present an application of the method to detect slow-moving Kuiper Belt Objects (KBOs) in the New Horizons target search observations. A cadence that takes continuous images of one HSC field of view for half a night fits the method well. The observations for the New Horizons Kuiper Belt Extended Mission (NH/KEM) using HSC began in May 2020, and are ongoing. Here we show our result of the analysis of the dataset acquired from May 2020 through June 2021 that have already passed the proprietary period and are open to the public. We detected 84 KBO candidates in the June 2020 and June 2021 datasets, when the observation field was close to opposition. △ Less

Submitted 8 July, 2024; originally announced July 2024.

Comments: 21 pages, 9 figures, 5 tables, accepted for publication on Publications of the Astronomical Society of Japan

arXiv:2407.05609 [pdf, other]

Open-world Multi-label Text Classification with Extremely Weak Supervision

Authors: Xintong Li, Jinya Jiang, Ria Dharmani, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang

Abstract: We study open-world multi-label text classification under extremely weak supervision (XWS), where the user only provides a brief description for classification objectives without any labels or ground-truth label space. Similar single-label XWS settings have been explored recently, however, these methods cannot be easily adapted for multi-label. We observe that (1) most documents have a dominant cl… ▽ More We study open-world multi-label text classification under extremely weak supervision (XWS), where the user only provides a brief description for classification objectives without any labels or ground-truth label space. Similar single-label XWS settings have been explored recently, however, these methods cannot be easily adapted for multi-label. We observe that (1) most documents have a dominant class covering the majority of content and (2) long-tail labels would appear in some documents as a dominant class. Therefore, we first utilize the user description to prompt a large language model (LLM) for dominant keyphrases of a subset of raw documents, and then construct a (initial) label space via clustering. We further apply a zero-shot multi-label classifier to locate the documents with small top predicted scores, so we can revisit their dominant keyphrases for more long-tail labels. We iterate this process to discover a comprehensive label space and construct a multi-label classifier as a novel method, X-MLClass. X-MLClass exhibits a remarkable increase in ground-truth label space coverage on various datasets, for example, a 40% improvement on the AAPD dataset over topic modeling and keyword extraction methods. Moreover, X-MLClass achieves the best end-to-end multi-label classification accuracy. △ Less

Submitted 8 July, 2024; originally announced July 2024.

Comments: Preprint

arXiv:2407.04486 [pdf, other]

Variational and Explanatory Neural Networks for Encoding Cancer Profiles and Predicting Drug Responses

Authors: Tianshu Feng, Rohan Gnanaolivu, Abolfazl Safikhani, Yuanhang Liu, Jun Jiang, Nicholas Chia, Alexander Partin, Priyanka Vasanthakumari, Yitan Zhu, Chen Wang

Abstract: Human cancers present a significant public health challenge and require the discovery of novel drugs through translational research. Transcriptomics profiling data that describes molecular activities in tumors and cancer cell lines are widely utilized for predicting anti-cancer drug responses. However, existing AI models face challenges due to noise in transcriptomics data and lack of biological i… ▽ More Human cancers present a significant public health challenge and require the discovery of novel drugs through translational research. Transcriptomics profiling data that describes molecular activities in tumors and cancer cell lines are widely utilized for predicting anti-cancer drug responses. However, existing AI models face challenges due to noise in transcriptomics data and lack of biological interpretability. To overcome these limitations, we introduce VETE (Variational and Explanatory Transcriptomics Encoder), a novel neural network framework that incorporates a variational component to mitigate noise effects and integrates traceable gene ontology into the neural network architecture for encoding cancer transcriptomics data. Key innovations include a local interpretability-guided method for identifying ontology paths, a visualization tool to elucidate biological mechanisms of drug responses, and the application of centralized large scale hyperparameter optimization. VETE demonstrated robust accuracy in cancer cell line classification and drug response prediction. Additionally, it provided traceable biological explanations for both tasks and offers insights into the mechanisms underlying its predictions. VETE bridges the gap between AI-driven predictions and biologically meaningful insights in cancer research, which represents a promising advancement in the field. △ Less

Submitted 5 July, 2024; originally announced July 2024.

arXiv:2407.02899 [pdf, other]

Measurement of the branching fraction of the decay $J/ψ\to p \bar{p} η$

Authors: BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson, O. Afedulidis, X. C. Ai, R. Aliberti, A. Amoroso, Q. An, Y. Bai, O. Bakina, I. Balossino, Y. Ban, H. -R. Bao, V. Batozskaya, K. Begzsuren, N. Berger, M. Berlowski, M. Bertani, D. Bettoni, F. Bianchi, E. Bianco, A. Bortone, I. Boyko, R. A. Briere , et al. (639 additional authors not shown)

Abstract: A high precision measurement of the branching fraction of the decay $J/ψ\to p \bar{p} η$ is performed using $(10 087 \pm 44) \times 10^6$ $J/ψ$ events recorded by the {BESIII} detector at the {BEPCII} storage ring. The branching fractions of the two decays $J/ψ\to p \bar{p} η(η\to γγ)$ and $J/ψ\to p \bar{p} η(η\to π^+ π^- π^0)$ are measured individually to be… ▽ More A high precision measurement of the branching fraction of the decay $J/ψ\to p \bar{p} η$ is performed using $(10 087 \pm 44) \times 10^6$ $J/ψ$ events recorded by the {BESIII} detector at the {BEPCII} storage ring. The branching fractions of the two decays $J/ψ\to p \bar{p} η(η\to γγ)$ and $J/ψ\to p \bar{p} η(η\to π^+ π^- π^0)$ are measured individually to be $\mathcal{B}(J/ψ\to p \bar{p} η(η\to γγ)) = (1.480 \pm 0.001 \pm 0.024)\times\,10^{-3}$ and $\mathcal{B}(J/ψ\to p \bar{p} η(η\to π^+ π^- π^0)) = (1.557 \pm 0.003 \pm 0.038)\times\,10^{-3}$, where the first uncertainties are statistical and the second systematic. Both results are compatible within their uncorrelated systematic uncertainties. The combined result is $\mathcal{B}(J/ψ\to p \bar{p} η)=(1.495 \pm 0.001 \pm 0.023)\times\,10^{-3}$ where the first uncertainty is the combined statistical uncertainty and the second one the combined systematic uncertainty of both analyses, incorporating correlations between them. In addition, the $p \bar{p}$ threshold region is investigated for a potential threshold enhancement, and no evidence for one is observed. △ Less

Submitted 3 July, 2024; originally announced July 2024.

arXiv:2407.01601 [pdf, other]

Unveiling and Controlling Anomalous Attention Distribution in Transformers

Authors: Ruiqing Yan, Xingbo Du, Haoyu Deng, Linghan Zheng, Qiuzhuang Sun, Jifang Hu, Yuhang Shao, Penghao Jiang, Jinrong Jiang, Lian Zhao

Abstract: With the advent of large models based on the Transformer architecture, researchers have observed an anomalous phenomenon in the Attention mechanism--there is a very high attention on the first element, which is prevalent across Transformer-based models. It is crucial to understand it for the development of techniques focusing on attention distribution, such as Key-Value (KV) Cache compression and… ▽ More With the advent of large models based on the Transformer architecture, researchers have observed an anomalous phenomenon in the Attention mechanism--there is a very high attention on the first element, which is prevalent across Transformer-based models. It is crucial to understand it for the development of techniques focusing on attention distribution, such as Key-Value (KV) Cache compression and infinite extrapolation; however, the latent cause leaves to be unknown. In this paper, we analyze such a phenomenon from the perspective of waiver phenomenon, which involves reducing the internal values of certain elements in the sequence, allowing them to absorb excess attention without affecting their contribution to information. In specific models, due to differences in positional encoding and attention patterns, we have found that the selection of waiver elements by the model can be categorized into two methods: positional-encoding-based and feature-distribution-within-elements-based. △ Less

Submitted 3 July, 2024; v1 submitted 26 June, 2024; originally announced July 2024.

arXiv:2407.01492 [pdf, other]

RegMix: Data Mixture as Regression for Language Model Pre-training

Authors: Qian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng, Longxu Dou, Tianyu Pang, Jing Jiang, Min Lin

Abstract: The data mixture for large language model pre-training significantly impacts performance, yet how to determine an effective mixture remains unclear. We propose RegMix to automatically identify a high-performing data mixture by formulating it as a regression task. RegMix involves training a set of small models with diverse data mixtures and fitting a regression model to predict their performance gi… ▽ More The data mixture for large language model pre-training significantly impacts performance, yet how to determine an effective mixture remains unclear. We propose RegMix to automatically identify a high-performing data mixture by formulating it as a regression task. RegMix involves training a set of small models with diverse data mixtures and fitting a regression model to predict their performance given their respective mixtures. With the fitted regression model, we simulate the top-ranked mixture and use it to train a large-scale model with orders of magnitude more compute. To empirically validate RegMix, we train 512 models with 1M parameters for 1B tokens of different mixtures to fit the regression model and find the optimal mixture. Using this mixture we train a 1B parameter model for 25B tokens (i.e. 1000x larger and 25x longer) which we find performs best among 64 candidate 1B parameter models with other mixtures. Further, our method demonstrates superior performance compared to human selection and achieves results that match or surpass DoReMi, while utilizing only 10% of the compute budget. Our experiments also show that (1) Data mixtures significantly impact performance with single-task performance variations of up to 14.6%; (2) Web corpora rather than data perceived as high-quality like Wikipedia have the strongest positive correlation with downstream performance; (3) Domains interact in complex ways often contradicting common sense, thus automatic approaches like RegMix are needed; (4) Data mixture effects transcend scaling laws, and our approach captures the complexity by considering all domains together. Our code is available at https://github.com/sail-sg/regmix. △ Less

Submitted 1 July, 2024; originally announced July 2024.

arXiv:2407.01080 [pdf, other]

doi 10.1145/3637528.3671656

Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese

Authors: Yunqi Xu, Tianchi Cai, Jiyan Jiang, Xierui Song

Abstract: The prevailing issue of factual inconsistency errors in conventional Retrieval Augmented Generation (RAG) motivates the study of Factual Consistency Evaluation (FCE). Despite the various FCE methods proposed earlier, these methods are evaluated on datasets generated by specific Large Language Models (LLMs). Without a comprehensive benchmark, it remains unexplored how these FCE methods perform on o… ▽ More The prevailing issue of factual inconsistency errors in conventional Retrieval Augmented Generation (RAG) motivates the study of Factual Consistency Evaluation (FCE). Despite the various FCE methods proposed earlier, these methods are evaluated on datasets generated by specific Large Language Models (LLMs). Without a comprehensive benchmark, it remains unexplored how these FCE methods perform on other LLMs with different error distributions or even unseen error types, as these methods may fail to detect the error types generated by other LLMs. To fill this gap, in this paper, we propose the first comprehensive FCE benchmark \emph{Face4RAG} for RAG independent of the underlying LLM. Our benchmark consists of a synthetic dataset built upon a carefully designed typology for factuality inconsistency error and a real-world dataset constructed from six commonly used LLMs, enabling evaluation of FCE methods on specific error types or real-world error distributions. On the proposed benchmark, we discover the failure of existing FCE methods to detect the logical fallacy, which refers to a mismatch of logic structures between the answer and the retrieved reference. To fix this issue, we further propose a new method called \emph{L-Face4RAG} with two novel designs of logic-preserving answer decomposition and fact-logic FCE. Extensive experiments show L-Face4RAG substantially outperforms previous methods for factual inconsistency detection on a wide range of tasks, notably beyond the RAG task from which it is originally motivated. Both the benchmark and our proposed method are publicly available.\footnote{\url{https://huggingface.co/datasets/yq27/Face4RAG}\label{link_face4rag}} △ Less

Submitted 3 July, 2024; v1 submitted 1 July, 2024; originally announced July 2024.

Journal ref: KDD 2024 (oral)

arXiv:2407.00024 [pdf, other]

LMVD: A Large-Scale Multimodal Vlog Dataset for Depression Detection in the Wild

Authors: Lang He, Kai Chen, Junnan Zhao, Yimeng Wang, Ercheng Pei, Haifeng Chen, Jiewei Jiang, Shiqing Zhang, Jie Zhang, Zhongmin Wang, Tao He, Prayag Tiwari

Abstract: Depression can significantly impact many aspects of an individual's life, including their personal and social functioning, academic and work performance, and overall quality of life. Many researchers within the field of affective computing are adopting deep learning technology to explore potential patterns related to the detection of depression. However, because of subjects' privacy protection con… ▽ More Depression can significantly impact many aspects of an individual's life, including their personal and social functioning, academic and work performance, and overall quality of life. Many researchers within the field of affective computing are adopting deep learning technology to explore potential patterns related to the detection of depression. However, because of subjects' privacy protection concerns, that data in this area is still scarce, presenting a challenge for the deep discriminative models used in detecting depression. To navigate these obstacles, a large-scale multimodal vlog dataset (LMVD), for depression recognition in the wild is built. In LMVD, which has 1823 samples with 214 hours of the 1475 participants captured from four multimedia platforms (Sina Weibo, Bilibili, Tiktok, and YouTube). A novel architecture termed MDDformer to learn the non-verbal behaviors of individuals is proposed. Extensive validations are performed on the LMVD dataset, demonstrating superior performance for depression detection. We anticipate that the LMVD will contribute a valuable function to the depression detection community. The data and code will released at the link: https://github.com/helang818/LMVD/. △ Less

Submitted 8 May, 2024; originally announced July 2024.

arXiv:2406.19915 [pdf, other]

Correlated insulators in twisted mono-mono-bilayer graphene

Authors: Jin Jiang, Kenji Watanabe, Takashi Taniguchi, Mitali Banerjee

Abstract: Graphene-based moiré superlattice featuring flat bands is a promising platform for studying strongly correlated states. By tuning two twist angles and displacement fields in twisted mono-mono-bilayer graphene (TMMBG), we observed a correlated metallic state evolved into a valley-polarized correlated insulator at the filling factor of $ν= -2$, with the coupling strength between the top two monolaye… ▽ More Graphene-based moiré superlattice featuring flat bands is a promising platform for studying strongly correlated states. By tuning two twist angles and displacement fields in twisted mono-mono-bilayer graphene (TMMBG), we observed a correlated metallic state evolved into a valley-polarized correlated insulator at the filling factor of $ν= -2$, with the coupling strength between the top two monolayer graphene intensifying. Moreover, the corresponding effective g-factor obtained from fitting the thermal activation gap is enhanced with the coupling strength, suggesting that TMMBG can be harnessed for developing valleytronics devices. In addition, the observation of an unconventional correlated electron-hole insulator suggests that this state might be a candidate for an excitonic insulator, which may generated by the correlation from band nesting, which encourages further research in non-Fermi liquid physics. Our work reveals that tuning multiple twist angles can provide a unique route for studying quantum many-body states in twistronics. △ Less

Submitted 28 June, 2024; originally announced June 2024.

arXiv:2406.19735 [pdf, other]

Pseudoscalar heavy quarkonium production in heavy ion ultraperipheral collision

Authors: Jun Jiang, Shi-Yuan Li, Xiao Liang, Yan-Rui Liu, Cong-Feng Qiao, Zong-Guo Si, Hao Yang

Abstract: The inclusive production of pseudoscalar heavy quarkoniua ($η_c,~η_b$ and $B_c$) via photon-photon fusion in heavy ion ultraperipheral collision (UPC) are calculated to QCD next-to-leading order in the framework of non-relativistic QCD (NRQCD). The total cross section of $η_c$ produced in Pb-Pb UPC is 194 $\mathrm{nb}^{-1}$ and 1052 $\mathrm{nb}^{-1}$ at nucleon-nucleon c.m. energies… ▽ More The inclusive production of pseudoscalar heavy quarkoniua ($η_c,~η_b$ and $B_c$) via photon-photon fusion in heavy ion ultraperipheral collision (UPC) are calculated to QCD next-to-leading order in the framework of non-relativistic QCD (NRQCD). The total cross section of $η_c$ produced in Pb-Pb UPC is 194 $\mathrm{nb}^{-1}$ and 1052 $\mathrm{nb}^{-1}$ at nucleon-nucleon c.m. energies $\sqrt{S_{\mathrm{NN}}}=$ 5.52 TeV and 39.4 TeV, respectively. The cross sections for $η_b$ and $B_c$ mesons are more than two to three orders of magnitude smaller. We make a detailed phenomenological analysis on the $η_c$ production; the uncertainties caused by the renormalization scale and the charm quark mass, the cross sections in other ultraperipheral nucleon-nucleon colliding systems, and the transverse momentum distribution are discussed. At the coming HL-LHC and future FCC, the heavy ion UPC opens another door of the study on the production of heavy quarkonium. △ Less

Submitted 28 June, 2024; originally announced June 2024.

Comments: 16 pages, 1 figures, 4 tables

arXiv:2406.19631 [pdf, other]

Personalized Interpretation on Federated Learning: A Virtual Concepts approach

Authors: Peng Yan, Guodong Long, Jing Jiang, Michael Blumenstein

Abstract: Tackling non-IID data is an open challenge in federated learning research. Existing FL methods, including robust FL and personalized FL, are designed to improve model performance without consideration of interpreting non-IID across clients. This paper aims to design a novel FL method to robust and interpret the non-IID data across clients. Specifically, we interpret each client's dataset as a mixt… ▽ More Tackling non-IID data is an open challenge in federated learning research. Existing FL methods, including robust FL and personalized FL, are designed to improve model performance without consideration of interpreting non-IID across clients. This paper aims to design a novel FL method to robust and interpret the non-IID data across clients. Specifically, we interpret each client's dataset as a mixture of conceptual vectors that each one represents an interpretable concept to end-users. These conceptual vectors could be pre-defined or refined in a human-in-the-loop process or be learnt via the optimization procedure of the federated learning system. In addition to the interpretability, the clarity of client-specific personalization could also be applied to enhance the robustness of the training process on FL system. The effectiveness of the proposed method have been validated on benchmark datasets. △ Less

Submitted 27 June, 2024; originally announced June 2024.

arXiv:2406.19190 [pdf, ps, other]

Improved measurement of the semileptonic decay $D^+_{s}\to K^0 e^+ν_e$

Authors: BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson, O. Afedulidis, X. C. Ai, R. Aliberti, A. Amoroso, Q. An, Y. Bai, O. Bakina, I. Balossino, Y. Ban, H. -R. Bao, V. Batozskaya, K. Begzsuren, N. Berger, M. Berlowski, M. Bertani, D. Bettoni, F. Bianchi, E. Bianco, A. Bortone, I. Boyko, R. A. Briere , et al. (643 additional authors not shown)

Abstract: Analyzing $e^+e^-$ collision data corresponding to an integrated luminosity of $7.33~\mathrm{fb}^{-1}$ collected at center-of-mass energies between 4.128 and 4.226~GeV with the BESIII detector, we measure the branching fraction of the semileptonic decay $D^+_{s}\to K^0 e^+ν_e$ to be $(2.98\pm0.23\pm0.12)\times10^{-3}$. The $D_s^+\to K^0$ hadronic form factor is determined from the differential dec… ▽ More Analyzing $e^+e^-$ collision data corresponding to an integrated luminosity of $7.33~\mathrm{fb}^{-1}$ collected at center-of-mass energies between 4.128 and 4.226~GeV with the BESIII detector, we measure the branching fraction of the semileptonic decay $D^+_{s}\to K^0 e^+ν_e$ to be $(2.98\pm0.23\pm0.12)\times10^{-3}$. The $D_s^+\to K^0$ hadronic form factor is determined from the differential decay rate of $D^+_s\to K^0 e^+ν_e$ to be $f^{K^0}_+(0)=0.636\pm0.049\pm0.013$. For both measurements, the first uncertainty is statistical and the second systematic. The branching fraction and form factor measurements are factors of 1.6 and 1.7 more precise than the previous world averages, respectively. △ Less

Submitted 27 June, 2024; originally announced June 2024.

Comments: 13 pages, 6 figures

arXiv:2406.19138 [pdf]

doi 10.1016/j.molliq.2024.125362

Crude Oil Displacement Enhanced by Interfacially Active Nanoparticles and Their Coupling Effect with Low-Salinity Brines

Authors: Suparit Tangparitkul, Anupong Sukee, Jiatong Jiang, David Harbottle

Abstract: From the microscopic scale to the petroleum-reservoir scale, the interfacial phenomena of the crude oil-water-rock system crucially control an immiscible flow in a porous reservoir. One of the key mechanisms is crude oil droplet displacement dynamics, which can be optimized by manipulating the oil-water interfacial tension and the three-phase contact angle by means of chemical injection. The curre… ▽ More From the microscopic scale to the petroleum-reservoir scale, the interfacial phenomena of the crude oil-water-rock system crucially control an immiscible flow in a porous reservoir. One of the key mechanisms is crude oil droplet displacement dynamics, which can be optimized by manipulating the oil-water interfacial tension and the three-phase contact angle by means of chemical injection. The current study primarily investigated oil displacement enhanced by interfacially active nanoparticles, namely poly(N-isopropylacrylamide) or pNIPAM, which found an acceleration of oil droplet receding rate (5.66 degree/s) and a greater degree of oil droplet dewetted (37.0 degree contact angle). This was due to a contribution from the nanoparticle-induced structural disjoining pressures between the oil-water and water-solid interfaces. The coupling effect of pNIPAM nanoparticles with low-salinity brines was examined, which revealed a discrepancy in different brine valences. Coupling with divalent CaCl2 led to much slower oil droplet receding dynamics (2.58 degree/s, and 58.7 degree contact angle) since oil-substrate bridging is formulated and promoted by the divalent cation. However, positive synergy was observed with a monovalent NaCl blend. The crude oil dewetting dynamics were enhanced (9.55 degree/s) owing to the combined salt-induced hydration and nanoparticle-induced structural forces. The contact angle was as low as 21.5 degree before eventually detaching from the substrate for a relatively short period (156 s). These findings highlight the coupling effect of nanoparticles and low-salinity brine on the dewetting of heavy crude oil. Adding nanoparticles to an optimal brine could be an option for faster and greater fluid displacement, which is not limited to oil production applications but several others, such as detergency and other forms of geological storage. △ Less

Submitted 27 June, 2024; originally announced June 2024.

Comments: 15 pages

arXiv:2406.18965 [pdf, ps, other]

Exotic 4f Correlated Electronic States of Ferromagnetic Kondo Lattice Compounds ReRh$_6$Ge$_4$ (Re=Ce, Ho, Er, Tm)

Authors: Yu Gao, Jun Jiang, Haiyan Lu, Qiaoni Chen

Abstract: CeRh$_6$Ge$_4$ stands out as the first stoichiometric metallic compound with a ferromagnetic quantum critical point, thereby garnering significant attention. Ferromagnetic Kondo lattice compounds ReRh$_6$Ge$_4$ (Re=Ce, Ho, Er, Tm) have been systematically investigated with density functional theory incorporating Coulomb interaction U and spin-orbital coupling. We determined the magnetic easy axis… ▽ More CeRh$_6$Ge$_4$ stands out as the first stoichiometric metallic compound with a ferromagnetic quantum critical point, thereby garnering significant attention. Ferromagnetic Kondo lattice compounds ReRh$_6$Ge$_4$ (Re=Ce, Ho, Er, Tm) have been systematically investigated with density functional theory incorporating Coulomb interaction U and spin-orbital coupling. We determined the magnetic easy axis of CeRh$_6$Ge$_4$ is within the ab plane, which is in agreement with previous magnetization measurements conducted under external magnetic field and muSR experiments. We also predicted the magnetic easy axes for the other three compounds. For TmRh$_6$Ge$_4$, the magnetic easy axis aligns along the c axis, thus preserving the $C_3$ rotational symmetry of the c axis. Especially, there are triply degenerate nodal points along the $Γ-A$ direction in the band structure including spin-orbital coupling. A possible localized to itinerant crossover is revealed as $4f$ electrons increase from CeRh$_6$Ge$_4$ to TmRh$_6$Ge$_4$. Specifically, the $4f$ electrons of TmRh$_6$Ge$_4$ contribute to the formation of a large Fermi surface, indicating their participation in the conduction process. Conversely, the $4f$ electrons in HoRh$_6$Ge$_4$, ErRh$_6$Ge$_4$ and CeRh$_6$Ge$_4$ remain localized, which result in smaller Fermi surfaces for these compounds. These theoretical investigations on electronic structure and magnetic properties shed deep insight into the unique nature of $4f$ electrons, providing critical predictions for subsequent experimental studies. △ Less

Submitted 27 June, 2024; originally announced June 2024.

arXiv:2406.18846 [pdf, other]

AFBench: A Large-scale Benchmark for Airfoil Design

Authors: Jian Liu, Jianyu Wu, Hairun Xie, Guoqing Zhang, Jing Wang, Wei Liu, Wanli Ouyang, Junjun Jiang, Xianming Liu, Shixiang Tang, Miao Zhang

Abstract: Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse design, which requires to generate and edit diverse geometric-qualified and aerodynamic-qualified ai… ▽ More Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse design, which requires to generate and edit diverse geometric-qualified and aerodynamic-qualified airfoils following the multimodal instructions, \emph{i.e.,} dragging points and physical parameters. This paper presents the open-source endeavors in airfoil inverse design, \emph{AFBench}, including a large-scale dataset with 200 thousand airfoils and high-quality aerodynamic and geometric labels, two novel and practical airfoil inverse design tasks, \emph{i.e.,} conditional generation on multimodal physical parameters, controllable editing, and comprehensive metrics to evaluate various existing airfoil inverse design methods. Our aim is to establish \emph{AFBench} as an ecosystem for training and evaluating airfoil inverse design methods, with a specific focus on data-driven controllable inverse design models by multimodal instructions capable of bridging the gap between ideas and execution, the academic research and industrial applications. We have provided baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained on an RTX 3090 GPU within 16 hours. The codebase, datasets and benchmarks will be available at \url{https://hitcslj.github.io/afbench/}. △ Less