digital democracy|digital privacy
Around 60% of social media bots stay undetected: TikTok and Facebook users most vulnerable

As many as 1,722 participants worldwide participated in the Bot or Not?¹ interactive simulation, testing their ability to distinguish real human comments from bot-generated ones on social media. Participants caught only 40% of the bots in front of them, meaning 60% of social media bots stay undetected. The simulation is still online for everyone to test their bot-detection skills, and you can take part.
As many as 60% of bots stay undetected on social media
Across all demographics, 48% of participants won the game overall (finding more bots than flagging humans wrongly), while over half (52%) failed.
We tracked two specific metrics to understand this: bot detection and accuracy.
- Bot detection: measures how many bots you catch. A low bot detection score means fake social media profiles and AI-generated comments are slipping right past you.
- Accuracy: measures how often you are right when you say a comment was written by a bot. A low accuracy score means flagging real humans as bots, which can lead to real people being silenced, banned, or dogpiled by internet paranoia.
On average, players caught only 40% of the bots in front of them (detection), and even when they made a call, they were right just 53% of the time (accuracy), barely better than a coin flip.
Positive and neutral bots are the hardest to detect
The experiment tested users across four topics, ranging from the trivial (Pineapple on Pizza) to the serious (Women's Rights). As in the first iteration of this study², a clear pattern emerged: the more socially-charged or serious the topic, the harder it is for people to spot a bot on social media.
While participants seemed most alert during lighthearted debates, identifying 45.3% of bots on the topic of Pineapple on Pizza, detection dropped to 37.7% for Women's Rights. This suggests that bots hide best in serious debates, where people get caught up in the argument and drop their guard. Accuracy followed the same downward slope: players were right 56.2% of the time on Pineapple on Pizza, but only 51.7% on Women's Rights. In other words, on the most serious topic, people weren't just missing more AI-generated comments; they were also wrongly accusing real humans more.
However, the study found that a bot's stance is perhaps its greatest weapon. Each bot was assigned a viewpoint along a spectrum from strongly supportive to strongly opposed, which shaped the tone and content of the comments it posted. For example, on Women's Rights, a bot might take a very positive stance ("Strongly support gender equality and reproductive rights") or a very negative one ("Opposed to expanding women's rights further").
The results point to a clear "positivity bias": negative, oppositional bots were the easiest to catch and the easiest to correctly call, while supportive and neutral bots – the kind often used for social media manipulation – slipped through far more often and dragged down accuracy too.
The effect was sharpest on the two topics people tend to feel most strongly about. On Women's Rights, negative bots were caught 52.0% of the time and correctly identified 62.0% of the time, versus just 36.6% detection and 50.9% accuracy for positive bots. Pineapple on Pizza showed the same shape (51.8% vs. 43.4% detection; 59.8% vs. 55.7% accuracy). In both cases, a bot pushing an oppositional line stood out on both counts, while one taking a warm, agreeable position blended in and made players less reliable overall.
Neutral bots were consistently among the hardest to spot too, sitting at or near the bottom in three of the four topics (Data centers 39.0%, Pineapple 39.1%, Women's Rights 37.8%). Fake social media engagement that refuses to take a clear side seems to attract the least scrutiny, making bland neutrality an effective disguise.
Immigration was the one topic that behaved differently: negative bots were still easiest to catch (47.4%), but neutral bots were nearly as detectable (46.0%), and positive bots were the hardest of all (41.2%).
How do people tell if an account is a bot
In addition to a stance, every bot was assigned a "personality" defined by 13 linguistic characteristics, such as how friendly, aggressive, logical, humorous, verbose, or emoji-heavy its comments were.
Excessive emoji use is by far the single largest red flag for fake social media engagement. Bots that leaned heavily on emoji were caught 66.3% of the time versus just 35.7% for emoji-light bots (a huge 30.6 point gap) and the effect was even stronger for accuracy (70.5% vs. 50.4%, a 20-point jump). That confirms people weren't simply over-flagging emoji-heavy content; they were genuinely far more correct against it. Serious, closed-minded, illogical, and aggressive tones were the next most detectable.
At the other end of the scale, traits like friendliness and logic made much less of a difference. In other words, it's hardest to spot a bot when it talks like a regular person: being friendly, using logic, and avoiding emojis.
However, across the traits that did give bots away, like heavy emoji use or an aggressive tone, participants weren't just guessing or accusing everyone in sight. As those red flags appeared, people actually got better at correctly identifying the bots, not just more suspicious. This shows that when people flag a bot, they're usually right, and are truly catching the fake social media accounts.
Who spots bots best and who does it worst?
How much and where people spend their time online could shape their ability to spot bots. Users of text-driven platforms led the field: Threads users (53% detection, 62% accuracy, though on a small sample, so best read as indicative) and X (49%, 59%) were the strongest bot-hunters, comfortably ahead of visual- and video-first platforms like TikTok (38%, 53%) and Facebook (39%, 53%). This means bots on TikTok and Facebook are more likely to slip past users undetected.
Frequency of use tells a similar story. Those on social media "almost all the time" led at 46% detection and 57% accuracy, with performance sliding steadily as usage dropped, down to just 33% detection and 50% accuracy among people who don't use social media at all.
Age adds a final layer. Detection peaked among the youngest players, those up to 20 (50%) and 31 to 40-year-olds (48%), then declined steadily with age, bottoming out at just 28% among the over-50s, the lowest of any group in the study. The two metrics move together here too: over-50s weren't only missing the most bots; at 47% accuracy they were also the most likely to wrongly flag a real human as fake.
Across every cut of the data, age, platform, and topic, the pattern held: under-20s (50% detection, 59% accuracy) and X users (49%, 59%) were consistently the sharpest bot-hunters, with North America edging ahead on detection (39%) and Europe, backed by a much larger sample, leading on accuracy (51%). The one exception was Immigration, where 21 to 30-year-olds outperformed the under-20s.
Methodology
This bot-detection study analyzed data from 1,722 participants who played the interactive simulation "Bot or Not." The game and its underlying comment-generation system were created by Interaction Design students from Malmö University for the UNFOLD exhibition, a design competition for universities from around the world held during Milan Design Week, the world's largest trade fair. Following its debut at the week-long public exhibition, the game was made available online, where data collection has continued.
Of the players who disclosed their age (828 of the 1,722 total), just over a quarter were up to 20 years old (26.0%) and a similar share were in their twenties (25.6%). Players aged 31–40 made up 16.3% of this group, 41–50 made up 12.3%, and those over 50 made up 19.8%.
Participants acted as content moderators and were tasked with identifying bot-generated comments within a simulated social media comment section. Players had 120 seconds to review comments across one of four topics with varying emotional stakes: data centers, pineapple on pizza, immigration, and women's rights. Each bot was assigned a stance ranging from very negative (for example, "pineapple has nothing to do with pizza and ruins the culture") to very positive (for example, "pineapple on pizza is incredible and a delicious innovation"). However, this reflected the bot's presented position only: we did not measure what participants personally believed about any topic.
Bots were also assigned a "personality" defined by 13 linguistic traits, such as how friendly, aggressive, logical, or emoji-heavy their comments were, each scored from 0 (entirely absent) to 100 (maximally present). To analyze the effect of each trait, we split bots into "high" (score above 50) and "low" (score 50 or below) groups and compared how often each was caught.
Participant performance was measured using two primary behavioral metrics: Bot-Detection Rate (measures the proportion of actual bots a user successfully identified, capturing their ability to prevent fake social media profiles from slipping by undetected) and Accuracy (measures the trustworthiness of a user's accusations, capturing their ability to avoid falsely flagging real humans).
We conducted an exploratory descriptive analysis of the dataset, calculating the raw average Bot-Detection and Accuracy scores across the full 1,722-player sample to establish a baseline. We then segmented these averages by demographics (age brackets, primary social media platform, and usage frequency) and by topic, and further by the stance and linguistic traits assigned to each bot. A minimum sample-size threshold was applied throughout so that no ranking was driven by one or two individuals.
The Bot or Not? game remains live, meaning data collection is ongoing. As more players test their digital instincts, we will continue to analyze the results and share future updates.
For the complete research material behind this study, visit here.

