The Data Labeling Marketplace | Where AI Builders and AI Trainers Connect to Build the Future | Find, hire, & securely pay data labelers for any annotation tool

Seattle, Washington
OpenTrain AI retweeted
At this point I feel like we understand pretty well what's going on with LLMs: - Outputs are roughly equivalent to kernel smoothing over positional embeddings (arxiv.org/pdf/1908.11775.pdf) - The learned computation model is *probably* bounded by RASP-L (arxiv.org/pdf/2310.16028.pdf) - LLMs learn structure primarily from human generated content (text, images) which is far more structured and predictable than the universe. - LLaMa3 shows us that the higher quality the annotations on the human generated content, the better the LLMs do (10million messages is a lot!) - Multi-turn labeling is currently very expensive and so likely driving the costs of the models. - Right now we're likely bottlenecked not on CPU, or size of data, but number and quality of annotations. So tl;dr. Great at predicting what a human would do or say by averaging in distribution data in the corpus. No emergent generality. Currently bottlenecked by high quality annotated data. @hausdorff_space did I miss anything?
51
144
1,041
301,905
OpenTrain AI retweeted
The perfect quote to describe LLMs can be found in a 1946 Jean Cocteau movie -- "Réfléchissez pour moi, je réfléchirai pour vous" (think for me, I will reflect for you). What you get from the model is always a reflection of the training data you put in -- itself a by-product of human thinking.
22
88
512
79,781
NEW FEATURE: AI Live Chat Interviews: Using GPT-4 to interview, score, and screen job applicants from data labelers/AI trainers. One-on-one interviews, at scale! How It Works: #GPT4
3
5
1,733
Learn more about OpenTrain's AI Interviewer here: opentrain.ai/news/ai-powered…
1
1,017
Traditionally, selecting who to hire required significant time for: -Viewing applicants' resume -Reading cover letters -Messaging follow up questions -THEN judge if they are a good fit for your specific job
2
3
714
With the power of LLMs, we can now do all of this for you behind-the-scenes and then present an AI Qualification Score. From there, you can get a clear view of who are the most qualified candidates.
1
5
768
OpenTrain AI retweeted
102
628
5,865
356,785
Grok-1 by @xai utilized "AI Tutors", or human subject matter experts to create custom training data & provide RLHF. We're making this easy for anybody to do this with OpenTrain.ai: The data labeling marketplace to find, hire, & pay training data experts for ANY data labeling software. Simply setup the data workflows on your data labeling tool of choice, & use OpenTrain to find the subject matter experts to work from that tool. Like UpWork, but for "AI Tutoring" Post jobs & collect applicants 100% free! #rlhf #llm #datalabeling #trainingdata
2
8
1,101
OpenTrain AI retweeted
🔥Excited to introduce LMSYS-Chat-1M, a large-scale dataset of 1M real-world conversations with 25 cutting-edge LLMs! This dataset, collected from chat.lmsys.org, offers insights into user interactions with LLMs and intriguing use cases. Link: huggingface.co/datasets/lmsy…
9
84
356
96,312
An in-depth look at RLHF by @natolambert from @huggingface. The need for high-quality, task-specific data in RLHF is crucial. With OpenTrainAI, you can find, hire, & pay the human experts essential for responsible and effective RLHF. Post your job today! #RLHF #MachineLearning
Reinforcement Learning from Human Feedback (RLHF) is gaining traction. This field aims to make AI more responsible by including human values and preferences. In this video, @natolambert, a research scientist and RLHF team lead at @huggingface explores its inner workings, applications and industry impact. RLHF has gained the spotlight in recent years. The growth of language models like Anthropic’s Claude and OpenAI's ChatGPT have increased interest in human-feedback integration. "There are some rumors that Open AI had two teams; one was doing RLHF and the other instruction fine-tuning. And the RLHF team kept getting more and more performance." Understanding RLHF The RLHF process has three main steps: Pre-training: Much like with GPT models, the journey starts with pre-training on a large corpus of data. This can range from text data, web scrapes, to specialized datasets. Reward Modeling: This is the RLHF counterpart of supervised fine-tuning in large language models. This stage involves creating a reward model that resonates with human values and preferences. RL Optimization: This stage parallels reward modeling and reinforcement learning in traditional AI models. The AI system fine-tunes itself based on the reward model, employing reinforcement learning algorithms for that extra layer of optimization. The Data Challenge Data collection and curation in RLHF closely resemble the challenges you'd encounter in large language model training. Datasets from organizations like OpenAI can serve as a useful foundation. However, the need for high-quality, task-specific data cannot be overstated. Implementing RLHF: A Practical Guide If you’re someone who loves getting hands-on with AI libraries like Hugging Face, implementing RLHF is right way to do. It’s essential to understand its limitations. Think about model stability, over-optimization, and exploration strategies, much like you would when prompt engineering. Ongoing Research and Next Steps While he suggests that some basics figured out, there are layers of complexity that still need to be unraveled: 1. New Benchmarks: How do we measure the effectiveness of RLHF? 2. Preference Modeling: How can the model be made to understand human preferences better? 3. Interpreting RLHF: Much like explainability in traditional models, how do we make RLHF more interpretable? 4. System-Wide Evaluation: Going beyond individual performance, how does RLHF affect an entire system? The Transformative Power of RLHF Whether you're an AI developer, a business analyst, or a marketer, RLHF promises to revolutionize your domain. Imagine customer service chatbots that understand human emotions better, or content generators that align more closely with human values. RLHF is an emerging field that focuses on enhancing machine learning models through human feedback. While it tackles important issues like bias and ethics, its broader goal is to improve system performance across various applications. Whether you're deeply invested in the ethics of AI or simply curious about advancements in machine learning, RLHF offers valuable insights. If you're interested in the next wave of AI development, this area is definitely one to watch.
2
703
👀
When StackOverflow is fully dead (due to long congenital illness, self-inflicted wounds, and the finishing blow from AI), where will AI labs get their training data? They can just buy it! Assuming 10k quality answers per week, at $250/answer, that's just $130M/yr. Even at multiples of this estimate, quite affordable for large AI labs and big tech companies who are already spending much more than this on data.
520
👀
As discussed on @stratechery this morning, "experts are the new GPUs" and spending on data for training and fine tuning is in vertical takeoff.
499
OpenTrain AI retweeted
This is the way to unlock the next trillion high-quality tokens, currently frozen in textbook pixels that are not LLM-ready. Nougat: an open-source OCR model that accurately scans books with heavy math/scientific notations. It's ages ahead of other open OCR options. Meta is doing extraordinary open-source AI, sometimes without as much fanfare as Llama. My first serious AI research project (back @Columbia, 2012) was to convert chemical engineering PDFs into NLP-ready corpus. I still remember the immense pain of Tesseract, a much older OCR system (github.com/tesseract-ocr/tes…). Now Nougat runs a powerful Swin Transformer backbone and blows the benchmarks out of the water. We're talking about double-digit improvements across all metrics. Now, textbooks are all we need for the next GPT! Website: facebookresearch.github.io/n… Open-source code: github.com/facebookresearch/… Paper "Nougat: Neural Optical Understanding for Academic Documents": arxiv.org/abs/2308.13418
117
730
3,853
1,089,158