Forget chatbot training. AI's next big data grab is about learning how humans work. – Business Insider

Welcome to the forefront of conversational AI as we explore the fascinating world of AI chatbots in our dedicated blog series. Discover the latest advancements, applications, and strategies that propel the evolution of chatbot technology. From enhancing customer interactions to streamlining business processes, these articles delve into the innovative ways artificial intelligence is shaping the landscape of automated conversational agents. Whether you’re a business owner, developer, or simply intrigued by the future of interactive technology, join us on this journey to unravel the transformative power and endless possibilities of AI chatbots.
The AI race is shifting away from teaching chatbots how to answer questions and toward teaching agents how to perform entire jobs.
After years spent training models on internet text and paying contractors to rate chatbot responses, tech companies are increasingly focused on something new: creating realistic digital workplaces where AI can practice coding, using business software, making decisions, and completing long-running tasks much like a human employee.
These digital training realms are called reinforcement learning environments, or RL environments for short. Three recent developments show how this is catching on across the industry.
Google is in talks to invest more than $1.5 billion in Mechanize, a startup that builds virtual work environments to train AI agents, Business Insider exclusively reported last week. The potential deal would bring Mechanize talent to Google while the tech giant licenses some of its technology, with the team working on model evaluation and development.
This is the first in an ongoing series on The AI Data Grab.
This week, Elon Musk told SpaceX employees the company will collect and use their workplace activity data to help improve Grok AI models. “It will be trained on you,” the CEO said.
Meanwhile, Meta has been trying to collect employees’ keystrokes, mouse movements, clicks, and other screen activity, saying the goal is to help AI learn how people actually use computers — from keyboard shortcuts to navigating workplace software.
Taken together, these moves point to a broader shift in how frontier AI systems are being developed.
Take a smarter break in your day – and see how far you get.
The first generation of large language models was built on two main data foundations: enormous amounts of text scraped from books and the internet, followed by human feedback from contractors who judged whether chatbot answers were good or bad.
That approach created remarkably capable conversational AI. But it has struggled to produce AI agents that can independently complete complex, multi-step work over hours or days. Increasingly, the industry’s answer is RL environments, realistic simulations where AI agents learn by doing rather than simply predicting the next word.
Scale AI, one of the biggest suppliers of training data, says top AI companies are moving away from relying solely on static datasets and human preference feedback toward simulated environments where agents can safely learn through trial and error.
“The way models are trained is evolving,” Chetan Rane, head of product for agents & RL environments at Scale AI, wrote in a recent blog post. “Frontier models increasingly need to learn through trial and error in realistic simulated environments rather than relying solely on static datasets or human preference feedback.”
Nearly half of the company’s new AI training projects now involve RL environments, which model realistic coding, computer use, and enterprise workflows, he noted.
Instead of asking whether a chatbot produced the right answer, these environments allow an AI to attempt an entire workflow, make mistakes, recover, and improve based on whether it successfully completed the task.
Mechanize is betting this will become one of the most important parts of AI development.
The startup was launched last year by AI researcher Tamay Besiroglu with an unusually ambitious goal: “full automation of the economy.” At the time, he was criticized for such a bold mission.
“We will achieve this by creating simulated environments and evaluations that capture the full scope of what people do at their jobs,” Besiroglu and his cofounders wrote in a blog announcing Mechanize. “The market potential here is absurdly large: workers in the US are paid around $18 trillion per year in aggregate. For the entire world, the number is over three times greater, around $60 trillion per year.”
The startup says today’s AI systems remain unreliable at long-running work because they lack realistic environments in which to learn.
Its first target is software engineering. Mechanize argues that future coding agents will first learn from examples of professional programmers before improving through reinforcement learning inside increasingly realistic software environments that capture the complexity of real engineering projects.
As those environments improve, the company believes the same approach can expand into all kinds of white-collar work.
Reward signals are a key part of these RL environments, he told Business Insider in an interview last year.
“You want to be able to tell the model you did the task correctly versus incorrectly,” Besiroglu explained. “Then you want to leverage that to reinforce the kind of patterns of behavior that resulted in it correctly performing the task.”
That thinking offers a possible explanation for why large technology companies are suddenly interested in collecting data on how their own employees actually work.
In a SpaceX all-hands meeting this week, Musk told staff that collecting their workplace activity data will help align Grok AI models with better with human goals and ideals.
SpaceX employees, who he described as “a collection of some of the very best humans on Earth,” could imbue future models with desirable values, he added.
“You will effectively be the parents of the AI. It will inherit your thoughts and ideas and beliefs, and I think that’s a good thing,” the SpaceX CEO said.
Meta’s internal announcement said its tracking software would help AI understand how people complete everyday computer tasks because agents need to be trained on real examples. The software captures inputs such as mouse movements, clicks, keystrokes, and screen context across approved workplace applications.
Other companies are pursuing similar ideas.
Uber recently said it has begun embedding top AI engineers inside departments including finance, legal, HR, marketing, procurement, and customer support to observe how employees work before redesigning those workflows around AI. The company calls the initiative “Agentic Pods.”
The result is that the AI industry’s newest race may no longer be about building smarter chatbots.
Instead, companies are competing to build detailed digital versions of real workplaces, where AI agents can practice the thousands of decisions, actions, and workflows that make up modern jobs.
If the industry’s biggest players are right, these virtual workplaces will become the training grounds for the next generation of AI.
Sign up for BI’s Tech Memo newsletter here. Reach out to me via email at abarr@businessinsider.com.
No comments right now, check back later.
Comments are unavailable right now.
Jump to
Every time publishes a story, you’ll get an alert straight to your inbox!
Look out for an alert in your inbox the next time publishes a story!
Every time a new story is published, you’ll get an alert straight to your inbox!
Look out for an alert in your inbox the next time a new story is published!

By clicking “Sign up”, you agree to receive emails from Business Insider. In addition, you accept Insider’s Terms of Service and Privacy Policy.

source

Scroll to Top