Our mission is to organize science by converting information into useful knowledge.

London, UK
Based in Brazil
Filter
Exclude
Time range
-
Minimum likes
Thank you everyone for trying the Galactica model demo. We appreciate the feedback we have received so far from the community, and have paused the demo for now. Our models are available for researchers who want to learn more about the work and reproduce results in the paper.
30
45
505
This is just the first step on our mission to organize science. And there is a lot more work to be done. We look forward to seeing what the open ML community builds with the model.
12
8
120
Despite not being trained on a general corpus, Galactica outperforms BLOOM and OPT-175B on BIG-bench. Galactica is also significantly less toxic than other language models based on evaluations.
2
4
86
We train for over four epochs and experience improving performance with use of repeated tokens. For the largest 120B model, we trained for four epochs without overfitting.
1
4
100
Galactica performs well on reasoning, outperforming Chinchilla on mathematical MMLU by 41.3% to 35.7%, and PaLM 540B on MATH with a score of 20.4% versus 8.8%.
2
8
108
We release our initial paper below. We train on a large scientific corpus of papers, reference material, knowledge bases and many other sources. Includes scientific text and also scientific modalities such as proteins, compounds and more. galactica.org/paper.pdf
2
14
178
We believe models should be open. To accelerate science, we open source all models including the 120 billion model with no friction. You can access them here. github.com/paperswithcode/ga…
5
51
484
🪐 Introducing Galactica. A large language model for science. Can summarize academic literature, solve math problems, generate Wiki articles, write scientific code, annotate molecules and proteins, and more. Explore and get weights: galactica.org
208
1,980
7,343
We have explored some of the latest progress, architectural improvements, and emerging new techniques for long-range modeling. We'll continue to keep track of the progress on long-range modeling and LRA. More threads like this coming soon! Follow @paperswithcode for more. 10/10
7
Besides transformers, other types of models have been tested on LRA. Some of the top performing models are attained by S4 variants which are based on state space models. A recent, improved S4 variant (Liquid-S4) attained competitive results with Mega (current SoTA). 9/10
1
1
8
SoTA in the LRA benchmark is achieved by Mega (Ma et al. 2022), a single-head gated attention mechanism equipped with exponential moving average to incorporate inductive bias of position-aware local dependencies into the position-agnostic attention mechanism. 8/10
1
1
13
There’s also interest in assessing speed and memory footprint of ML models. Tay et al. (2020) reported efficiency results for a set of models. Performer and Linear Transformer make a better trade-off in terms of speed and performance with reasonable memory usage. 7/10
1
1
7
From models that have been evaluated on LRA, there's no clear winner. Some models like Linear Transformer and Performer perform well on text classification but don't do so well on data that is hierarchically structured (e.g., ListOps task). 6/10
1
6
There exists many solutions that have been built and designed to perform well at long-range modeling. Some notable transformer-based models include BigBird, Performer, Longformer, Synthesizer, among others. Check performance of these models here: paperswithcode.com/sota/long… 5/10
1
7
The LRA benchmark includes examples of sequences that range from 1K to 16K tokens. It consists of different data types and modalities such as text, synthetic images, mathematical expressions and so on. (Figure below shows ListOPs task included in the LRA benchmark) 4/10
1
6
Long-range models are evaluated using standard benchmarks. A popular benchmark used for evaluating model quality on long-context scenarios is called Long-Range Arena (LRA). LRA assesses generalization power and computational efficiency of ML models. paperswithcode.com/dataset/l… 3/10
1
1
7
Lots of ML models today have trouble modeling long sequences. Interesting problems that involve natural language generation require effective long-range modeling to produce good results so this has become an important capability to evaluate on. 2/10
1
7
How well do machine learning models perform on long sequences? This is a question of high interest in ML research so let’s take a look at what we know so far? 1/10
1
15
148
As language models continue to improve in capabilities, we will continue to track and document the progress on BIG-bench and other emerging benchmarks on @paperswithcode. 6/6
1
10