Radar van Elk Solutions

Science News · Wetenschap

AI agents aren’t ready to replace humans in behavioral research

A recent study suggests that AI digital twins, designed to mimic individual behavior for social science research, are not yet accurate enough to replace human subjects. While showing some promise, these AI surrogates often distort views and perform only slightly better than chatbots with basic demographic information.

Replacing human subjects with AI surrogates, or digital twins, is an idea some social scientists are considering due to the costs, fatigue, and potential distress associated with human participants. However, a study published in Science Advances indicates that these AI twins may distort their surrogate's views, creating a "funhouse mirror" effect.

Researchers, including Olivier Toubia from Columbia Business School, developed AI twins by feeding extensive personal data from over 2,000 individuals into a large language model. These twins were then tested across 19 social science experiments, evaluating responses to various scenarios, including political donations and algorithmic hiring. Although the twins performed better than chance, they were incorrect about a quarter of the time, performing similarly to chatbots that only received demographic data.

Despite their overall inaccuracies, the digital twins did capture more variation in responses compared to LLMs with limited information. For instance, they could differentiate between individuals' self-reported levels of self-control, whereas simpler models might average these differences out. This ability to reflect potential group differences, even with errors, offers some value.

The study attributes the AI twins' limitations to several factors: their responses tended to be more homogenous and skewed towards demographic stereotypes. Their accuracy also correlated with participants' affluence and education levels. Furthermore, the twins exhibited biases, such as increased trust and reduced concern about technological threats, and appeared more rational than their human counterparts.

Hadi Hosseini, an AI researcher at Penn State University, noted a similar tendency for LLMs to distort human judgment towards excessive rationality. He suggests that more dynamic training methods, such as having AI shadow individuals throughout their day or engage in regular conversations, could improve dataset quality.

Toubia acknowledges that while the tested digital twins have limitations, they could still be useful in specific research contexts. For example, their indefatigable nature allows for detailed responses where tired humans might provide brief answers. Pretesting experiments with digital twins could also help refine study designs before engaging human participants. However, Toubia emphasizes the need for realistic expectations regarding the predictive power of synthetic data in capturing the full spectrum of human behavior.

AI-samenvatting op basis van de bron.

Science News