subscribe to arXiv mailings

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Authors: Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love , et al. (1092 additional authors not shown)

Abstract: In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February… ▽ More In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February version on the great majority of capabilities and benchmarks; (2) Gemini 1.5 Flash, a more lightweight variant designed for efficiency with minimal regression in quality. Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks. Studying the limits of Gemini 1.5's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens, a generational leap over existing models such as Claude 3.0 (200k) and GPT-4 Turbo (128k). Finally, we highlight real-world use cases, such as Gemini 1.5 collaborating with professionals on completing their tasks achieving 26 to 75% time savings across 10 different job categories, as well as surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person who learned from the same content. △ Less

Submitted 14 June, 2024; v1 submitted 8 March, 2024; originally announced March 2024.

arXiv:2102.06166 [pdf, other]

Testing Framework for Black-box AI Models

Authors: Aniya Aggarwal, Samiulla Shaikh, Sandeep Hans, Swastik Haldar, Rema Ananthanarayanan, Diptikalyan Saha

Abstract: With widespread adoption of AI models for important decision making, ensuring reliability of such models remains an important challenge. In this paper, we present an end-to-end generic framework for testing AI Models which performs automated test generation for different modalities such as text, tabular, and time-series data and across various properties such as accuracy, fairness, and robustness.… ▽ More With widespread adoption of AI models for important decision making, ensuring reliability of such models remains an important challenge. In this paper, we present an end-to-end generic framework for testing AI Models which performs automated test generation for different modalities such as text, tabular, and time-series data and across various properties such as accuracy, fairness, and robustness. Our tool has been used for testing industrial AI models and was very effective to uncover issues present in those models. Demo video link: https://youtu.be/984UCU17YZI △ Less

Submitted 11 February, 2021; originally announced February 2021.

Comments: 4 pages Demonstrations track paper accepted at ICSE 2021

arXiv:1811.04665 [pdf, other]

What is my data worth? From data properties to data value

Authors: Kalapriya Kannan, Rema Ananthanarayanan, Sameep Mehta

Abstract: Data today fuels both the economy and advances in machine learning and AI. All aspects of decision making, at the personal and enterprise level and in governments are increasingly data-driven. In this context, however, there are still some fundamental questions that remain unanswered with respect to data. \textit{What is meant by data value? How can it be quantified, in a general sense?}. The "val… ▽ More Data today fuels both the economy and advances in machine learning and AI. All aspects of decision making, at the personal and enterprise level and in governments are increasingly data-driven. In this context, however, there are still some fundamental questions that remain unanswered with respect to data. \textit{What is meant by data value? How can it be quantified, in a general sense?}. The "value" of data is not understood quantitatively until it is used in an application and output is evaluated, and hence currently it is not possible to assess the value of large amounts of data that companies hold, categorically. Further, there is overall consensus that good data is important for any analysis but there is no independent definition of what constitutes good data. In our paper we try to address these gaps in the valuation of data and present a framework for users who wish to assess the value of data in a categorical manner. Our approach is to view the data as composed of various attributes or characteristics, which we refer to as facets, and which in turn comprise many sub-facets. We define the notion of values that each sub-facet may take, and provide a seed scoring mechanism for the different values. The person assessing the data is required to fill in the values of the various sub-facets that are relevant for the data set under consideration, through a questionnaire that attempts to list them exhaustively. Based on the scores assigned for each set of values, the data set can now be quantified in terms of its properties. This provides a basis for the comparison of the relative merits of two or more data sets in a structured manner, independent of context. The presence of context adds additional information that improves the quantification of the data value. △ Less

Submitted 12 November, 2018; originally announced November 2018.

arXiv:1711.04971 [pdf, other]

DataVizard: Recommending Visual Presentations for Structured Data

Authors: Rema Ananthanarayanan, Pranay Kr. Lohia, Srikanta Bedathur

Abstract: Selecting the appropriate visual presentation of the data such that it preserves the semantics of the underlying data and at the same time provides an intuitive summary of the data is an important, often the final step of data analytics. Unfortunately, this is also a step involving significant human effort starting from selection of groups of columns in the structured results from analytics stages… ▽ More Selecting the appropriate visual presentation of the data such that it preserves the semantics of the underlying data and at the same time provides an intuitive summary of the data is an important, often the final step of data analytics. Unfortunately, this is also a step involving significant human effort starting from selection of groups of columns in the structured results from analytics stages, to the selection of right visualization by experimenting with various alternatives. In this paper, we describe our \emph{DataVizard} system aimed at reducing this overhead by automatically recommending the most appropriate visual presentation for the structured result. Specifically, we consider the following two scenarios: first, when one needs to visualize the results of a structured query such as SQL; and the second, when one has acquired a data table with an associated short description (e.g., tables from the Web). Using a corpus of real-world database queries (and their results) and a number of statistical tables crawled from the Web, we show that DataVizard is capable of recommending visual presentations with high accuracy. We also present the results of a user survey that we conducted in order to assess user views of the suitability of the presented charts vis-a-vis the plain text captions of the data. △ Less

Submitted 14 November, 2017; originally announced November 2017.

Showing 1–4 of 4 results for author: Ananthanarayanan, R