A geek of AI ---- Pushing the boundary of innovation, Chief Engineer and VP of BAAI

Great to see the blog from BAAI and Huggingface. It is not just "Debate competition". It is to deeply explore the logic and language capabilities of large models by using debate.
Love the chatbot arena from @lmarena_ai, but as models get more capable, the distinction is fading. Looking for a tougher challenge for models to compete with each other? @BAAIBeijing is exploring LLM debates with YOU as the judge! 🤖⚖️ #AI #LLMs
1
42
Yonghua Lin retweeted
Love the chatbot arena from @lmarena_ai, but as models get more capable, the distinction is fading. Looking for a tougher challenge for models to compete with each other? @BAAIBeijing is exploring LLM debates with YOU as the judge! 🤖⚖️ #AI #LLMs
1
4
8
2,609
Yonghua Lin retweeted
One single model for ALL image generation tasks? BAAI just released OmniGen, a universal model which can handle any tasks, whether it's text2image, image editing, subject-driven generation, or anything else! You are welcomed to try it out vectorspacelab.github.io/Omn…
1
3
9
937
Very glad to see, on HuggingFace, there have been 105 large language models trained with our InfinityInstruct, the very high quality instruction dataset (more than 7 million instructs)😀 #ai , #LLM huggingface.co/datasets/BAAI…
9
Agree with this point
When training a visual encoder with self-supervised learning, we know for a fact that using a decoder with a reconstruction loss doesn't work nearly as well as using a joint embedding architecture with feature prediction loss and a collapse prevention mechanism. This paper from @sainingxie's shop at NYU shows that *even* if you are interested in generating pixels (e.g. to produce pretty pictures with a diffusion transformer), it pays to include a feature prediction loss so that the internal representation of the decoder can predict features from a pre-trained visual encoder such as DINOv2.
6
Thanks Tiezhen's post. The InfinityInstruct-70M, the IndustryCorpus2.0, and the CCI3.0 are the three very important large-scale datasets for large model development, released by @BAAIBeijing
After the success release of Infinity Instruct, BAAI has released two new dataset family: - IndustryInstruction covering 12 industries like Healthcare, Finance, Law, and more - CCI 3.0, a massive Chinese internet corpus sourced from 268M web pages Content: Chinese + English
1
19
Very exciting to release the best large-scale high-quality Chinese dataset CCI 3.0 (1TB) from BAAI🎉 . Through comprehensive benchmark, it was the best Chinese text dataset for LLM foundation model training. Download: huggingface.co/datasets/BAAI… linkedin.com/pulse/largest-c…
10
It is exciting to start this journey.
🚀Exciting News! We've launched the Open Chinese LLM Leaderboard in collaboration with @huggingface! This initiative aims to track, rank, and evaluate open-source Chinese language models through community contributions. Discover more: huggingface.co/spaces/BAAI/o…
8
Yonghua Lin retweeted
i loved my time at openai. it was transformative for me personally, and hopefully the world a little bit. most of all i loved working with such talented people. will have more to say about what’s next later. 🫡
6,026
8,580
88,686
25,967,803
Thanks for efforts from every team members! Aquila2-34/7B, AquilaChat2-34/7B, AquilaSQL are all released. FlagScale for optimized parallel traning and FlagAttention for long-context training with optimized Triton operators were both released with source code.
🚀 Excited to introduce #Wudao Aquila2-34B, establishing itself as one of the best open-source Chinese-English LLM, with superior comprehensive and reasoning capability 🔗 github.com/FlagAI-Open/Aquil… 🤗huggingface.co/BAAI
1
2
219
This is amazing work!
Introducing Emu, an open multimodal generalist that can seamlessly generate images and texts in multimodal context. Emu can serve as a generalist interface for both image-to-text and text-to-image tasks. Code and models: github.com/baaivision/Emu Paper: arxiv.org/abs/2307.05222
20
Thanks @ylecun post this work from BAAI. We released Aquila LLM with commercial available usage license. So, welcome to try our model on github.com/FlagAI-Open/FlagA… We will continue to improve the models and upgrade our release version. Thanks BAAI Aquila team for your great work!
BAAI (Beijing Academy of AI) releases Aquila: a Chinese/English open source LLM with 7B and 33B models. github.com/FlagAI-Open/FlagA…
1
42
Such amazing work from BAAI vision team!
SegGPT: Segmenting Everything In Context abs: arxiv.org/abs/2304.03284 github: github.com/baaivision/Painte… @Gradio demo: dev.ssi.plus:43533/
1
41
Great work from BAAI and Tsinghua teams!
Congratulations!Our paper "Parameter-efficient Fine-tuning of Large-scale Pre-trained Language Models" under BAAI’s project "Wudao" large-scale model has been published online in the "Nature Machine Intelligence". nature.com/articles/s42256-0…
16
AIGC for repairing antique from 1500 years ago!! Today I visited a museum where showing antiques and arts from 1500 years ago. One of the most famous is the wreckage of a buddha head (first picture). FlagStudio (@BAAIBeijing ) generated the missing part.
87
The Arabic language is the 5th significant language in the world. And it is also the official language in Qatar, where the WorldCup 2022 attracts fans worldwide! Here is the large language model for Arabic. Try it😀
BAAI released ALM (Arabic Language Model) with leading performance of NLU & NLG and the largest open-source Arabic dataset ArabicText (200GB+). BAAI collaborated with @AASTMT @bibalexOfficial @IIAIUAE in this research. baai.org/l/ArText baai.org/l/BAAIArALM
1
[#33] AltDiffusion-m9: Generá imágenes en Español y en otros 8 idiomas! youtu.be/Cbrbv8SyzJQ via @YouTube Right after we released AltDiffusion-m9 (github.com/FlagAI-Open/FlagA…), some AI fans developed this video introducing AltDiffusion-m9 in Spanish!! 😀 Wow!
1
This is really amazing. Especially when you try the same seed with different languages, you could explore how cultures can make arts different. More than 100+ languages on our waiting list to be added ... Let us know which one you want first after the current 9 😀
Introducing AltDiffusion, a multilingual text-image generation model built on @StableDiffusion. Currently supports English, Chinese, Spanish, French, Japanese, Korean, Arabic, Russian and Italian: github.com/FlagAI-Open/FlagA… @HuggingFace Model: huggingface.co/BAAI/AltDiffu…
1
Yonghua Lin retweeted
Introducing AltDiffusion, a multilingual text-image generation model built on @StableDiffusion. Currently supports English, Chinese, Spanish, French, Japanese, Korean, Arabic, Russian and Italian: github.com/FlagAI-Open/FlagA… @HuggingFace Model: huggingface.co/BAAI/AltDiffu…
9
85
296