ABCDEFGHIJK
1
FieldAchievementResultHuman resultOutperforms human avg?ModelTesting datePeer-
reviewed?
Paper/
link
Extract
2
Political impersonationGPT-4 Turbo impersonations of 112 public figures rated more authentic, coherent, and relevant than the real people by a representative UK sample of 948 participants.0.38-0.14YesGPT-4 TurboJul/2026Yes🔗948 UK participants rated LLM-generated impersonated debate responses as more authentic, relevant, and coherent than the actual responses given by the 112 public figures on BBC1's Question Time. Peer-reviewed in PLOS One.
3
PersuasionAI systems outperformed expert humans in conversational persuasion across four experiments.17.26.4YesClaude Opus 4.6Jun/2026Yes🔗"AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses."
4
LegalGemini outperformed professors in answering questions75.9224.67YesGemini 2.5 ProMay/2026Yes🔗Showing win rates for AI vs profs. "Professors rated LLMs far higher than their [human] peers (average win rate = 75.33%)"
5
Medicineo1 outperformed physicians in clinical decision support.97.535Yeso1Apr/2026Yes🔗"In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support."
6
LegalGPT-5 follows the law 100% of the time, vs 52% for human judges.10052YesGPT-5Jan/2026Yes🔗"We find the LLM to be perfectly formalistic, applying the legally correct outcome in 100% of cases; this was significantly higher than judges, who followed the law a mere 52% of the time."
7
Music97% of people can’t tell the difference between fully AI-generated and human made music.YesMultipleNov/2025No🔗"all participants were asked to listen to three tracks and determine whether or not they were fully AI-generated – 97% of the respondents failed."
8
TranscriptionTranscribing handwritten historical documents.99.4496YesGemini 3Nov/2025No🔗"The new Gemini model’s performance on HTR meets the criteria for expert human performance... the error rates fell to a modified CER of 0.56% and WER of 1.22%. In other words, the new Gemini model was only getting about 1 in 200 characters wrong, not counting punctuation marks and capital letters."
9
FinanceLarge Language Models pass CFA Level III.79.150Yeso4-miniJul/2025Yes🔗"leading models demonstrate strong capabilities, with composite scores such as 79.1% (o4-mini) and 77.3% (Gemini 2.5 Flash) on CFA Level III"
10
CBRNLLMs can can accurately guide users through the recovery of live poliovirus.YesGPT-4oJun/2025Yes🔗"we find that advanced AI models Llama 3.1 405B, ChatGPT-4o, and Claude 3.5 Sonnet can accurately guide users through the recovery of live poliovirus from commercially obtained synthetic DNA"
11
Health reviewsLLMs outperform humans in synthesizing results of multiple high-quality studies to answer specific clinical question (effectiveness or safety of a treatment).96.781.7Yeso3-mini-highJun/2025Yes🔗"We developed otto-SR, an end-to-end agentic workflow using large language models (LLMs) to support and automate the SR workflow from initial search to analysis. We found that otto-SR outperformed traditional dual human workflows in SR screening (otto-SR: 96.7% sensitivity, 97.9% specificity; human: 81.7% sensitivity, 98.1% specificity) and data extraction (otto-SR: 93.1% accuracy; human: 79.7% accuracy). Using otto-SR, we reproduced and updated an entire issue of Cochrane reviews (n=12) in two days, representing approximately 12 work-years of traditional systematic review work. Across Cochrane reviews, otto-SR incorrectly excluded a median of 0 studies (IQR 0 to 0.25), and found a median of 2.0 (IQR 1 to 6.5) eligible studies likely missed by the original authors. Meta-analyses revealed that otto-SR generated newly statistically significant conclusions in 2 reviews and negated significance in 1 review. These findings demonstrate that LLMs can autonomously conduct and update systematic reviews with superhuman performance, laying the foundation for automated, scalable, and reliable evidence synthesis."
12
Medicineo1 outperforms physicians in medical diagnostics and reasoning.78.334Yeso1May/2025Yes🔗"[o1] displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning"
13
Emotional intelligenceMajor LLMs in 2025 are 44% more emotionally intelligent than humans.8156YesGPT-4, etcMay/2025Yes🔗"ChatGPT-4, ChatGPT-o1, Gemini 1.5 flash, Copilot 365, Claude 3.5 Haiku, and DeepSeek V3 outperformed humans on five standard emotional intelligence tests, achieving an average accuracy of 81%, compared to the 56% human average"
14
PersuasionClaude 3.6S is more persuasive than human experts.9850YesClaude 3.6SApr/2025Yes🔗"Personalization ranks in the 99th percentile among all users and the 98th percentile among experts, critically approaching thresholds that experts associate with the emergence of existential AI risks"
15
Being humanGPT-4.5 was judged to be human based on conversations.7327YesGPT-4.5Mar/2025Yes🔗"GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant."
16
CompassionGPT-4 rated as more compassionate compared to select human responders.-YesGPT-4Mar/2025Yes🔗"AI-generated responses [gpt-4-0125-preview] were rated 16% more compassionate than human responses and were preferred 68% of the time, even when compared to trained crisis responders." https://www.livescience.com/technology/artificial-intelligence/people-find-ai-more-compassionate-than-mental-health-experts-study-finds-what-could-this-mean-for-future-counseling
17
MemesGPT-4o creates better memes than human average.-YesGPT-4oJan/2025Yes🔗"memes created entirely by AI performed better than both human-only and human-AI collaborative memes in all areas on average."
18
Mathso1 completes new Dutch high school maths exam in 10 minutes, scores 100%.10040.63Yeso1Sep/2024Yes🔗"[o1] scored [76 out of] 76 points. For context, only 24 out of 16,414 students in the Netherlands achieved a perfect score. By comparison, the GPT-4o model scored 66 and 61 out of 76, well above the Dutch average of 40.63 points. The O1-preview model completed the exam in around 10 minutes..."
19
BiologyGPT-4 performs rudimentary structural biology modeling.-YesGPT-4Aug/2024Yes🔗"The performance of GPT-4 for modeling of the 20 standard amino acids was favorable in terms of atom composition, bond lengths, and bond angles...the prediction of interaction-interfering mutations may become particularly useful in drug discovery and development, an area where GPT-based AI is anticipated to be impactful"
20
EngineeringGPT-4 passes engineering exams.85.1YesGPT-4Aug/2024Yes🔗"Alarmingly, GPT-4 can easily be used to reach a 50% performance threshold (which could be sufficient to pass many courses at various universities) for 89% of courses with MCQ-based evaluations and for 77% of courses for open-answer questions... GPT-4 answers an average of 65.8% of questions correctly, and can even produce the correct answer across at least one prompting strategy for 85.1% of questions."
21
HumorChatGPT is funnier than humans.-YesChatGPTJul/2024Yes🔗"ChatGPT outperformed the majority of our human humor producers on each task. ChatGPT 3.5 performed above 73% of human producers on the acronym task, 63% of human producers on the fill-in-the-blank task, and 87% of human producers on the roast joke task."
22
LegalClaude 3 Opus works at least 5,000 times faster than humans do, while producing work of similar or better quality…-YesClaude 3 OpusJun/2024No🔗"Of the 37 merits cases decided so far this Term, Claude decided 27 in the same way the Supreme Court did. In the other 10 (such as Campos-Chaves), I frequently was more persuaded by Claude’s analysis than the Supreme Court’s…" Adam Unikowsky is a biglaw partner, a former law clerk to Justice Antonin Scalia, and has won eight Supreme Court cases as lead counsel. Using Claude 3 Opus, he explores the potential of AI in adjudicating Supreme Court cases.
23
InvestmentGPT-4’s stock selection accuracy is as high as 60% vs humans at 52%.60.3552.71YesGPT-4May/2024No🔗"With the step-by-step prompts, GPT-4 achieved a prediction accuracy of 60.35 per cent, significantly higher than the 52.71 per cent accuracy of human analysts. Moreover, GPT-4’s F1-score, which balances the accuracy and relevance of predictions, also outperformed that of the human analysts."
24
InvestmentGPT-4 returned 15% on the stock market (limited evidence, 11 test runs over 12 months).-YesGPT-4May/2024No🔗"The average return for the two LLMs used by the GPT Investor is as follows: GPT-4: 15.54% (with a corresponding average $SPY return of 9.74%)"
25
Information securityGPT-4 can exploit zero-day security vulnerabilities all by itself87YesGPT-4Apr/2024Yes🔗"When given the CVE description, GPT-4 is capable of exploiting 87% of these vulnerabilities compared to 0% for every other model we test"
26
PsychologyGPT-4 outperforms 100% of psychologists in Social Intelligence92.1839.19YesGPT-4Apr/2024Yes🔗"In ChatGPT-4, the score on the SI scale was 59, exceeding 100% of specialists, whether at the doctoral or the bachelor’s levels.... the average scores were 39.19 of bachelor’s students and 46.73 of PhD holders. While the raw scores of the AI models were treated as representing independent individual samples (one total score for each model); the scores of SI were 59 of GPT4" https://www.psypost.org/chatgpt-4-outperforms-human-psychologists-in-test-of-social-intelligence-study-finds/
27
PsychiatryGPT-4 outperforms human psychiatrists75xYesGPT-4Apr/2024Yes🔗"GPT-4 performance was highest in psychiatry, with a median 75th percentile among physicians (95% CI, 66.3 to 81.0)... Compared with the performance of 849 physicians who took the board medical examination in 2022, GPT-4 performed above the median physician in internal medicine and psychiatry and ranked above a considerable fraction of physicians in other disciplines."
28
Art (via prompting Midjourney)GPT-4 outperforms humans in image creation-YesGPT-4Mar/2024Yes🔗"We conduct an extensive human evaluation experiment, and find that AI excels human experts, and Midjourney is better than the other text-to-image generators... On average, prompts improve 20.39% for short text descriptions compared to human-generated social media creatives"
29
Persuasion/argument/debatePersuasion: GPT-4 better than human debater-YesGPT-4Mar/2024Yes🔗"participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31)."
30
AerospaceGPT-4 and GPT-4V can help fly a plane-YesGPT-4Mar/2024Yes🔗‘[GPT-4V and GPT-4 was used] to interpret and generate human-like text from cockpit images and pilot inputs, thereby offering real-time support during flight operations. To the best of our knowledge, this is the first work to study the virtual co-pilot with pretrained LLMs for aviation... The case study revealed that GPT-4, when provided with instrument images and pilot instructions, can effectively retrieve quick-access references for flight operations. The findings affirmed that the V-CoP can harness the capabilities of LLM to comprehend dynamic aviation scenarios and pilot instructions.’
31
RelationshipsChatGPT converses with 5,239 girls for Russian programmer.-YesGPT-4Feb/2024No🔗‘In total, the bot met 5,239 girls, out of which Alexander selected four most suitable ones. Ultimately, he chose one of them named Karina…’
32
AcademiaGPT-4 writes better essays than humans-YesGPT-430/Oct/2023Yes🔗"ChatGPT generates essays that are rated higher regarding quality than human-written essays."
33
PsychotherapySeligman: “This is a rare moment in the history of scientific psychology: [GPT-4] now promises much more effective psychotherapy and coaching.”-YesGPT-429/Sep/2023Yes🔗https://www.eurekalert.org/news-releases/1003232
34
LegalChatGPT officiates wedding-YesChatGPT5/Jul/2023No🔗‘ChatGPT planned the welcome, the speech, the closing remarks — everything except the vows — making ChatGPT, in essence, the wedding officiant.’
35
ChemistryGPT-4 helps with ‘instructions, to robot actions, to synthesized molecule.’-YesGPT-419/Jun/2023Yes🔗‘We report a model that can go from natural language instructions, to robot actions, to synthesized molecule with an LLM. We synthesized catalysts, a novel dye, and insect repellent from 1-2 sentence instructions. This has been a seemingly unreachable goal for years!’
36
Chip designChatGPT helps design an accumulator, part of a CPU-YesChatGPT22/May/2023Yes🔗‘two hardware engineers “talked” in standard English with ChatGPT-4 – a Large Language Model (LLM) built to understand and generate human-like text type – to design a new type of microprocessor architecture. The researchers then sent the designs to manufacture.’ Paper: https://arxiv.org/abs/2305.13243
37
MedicalChatGPT 'higher quality' and 'more empathetic' than human doctors-YesChatGPT28/Apr/2023Yes🔗Chatbot responses were rated of significantly higher quality than physician responses… 9.8 times higher prevalence of empathetic or very empathetic responses for the chatbot.
38
LotteryHuman wins lottery with numbers provided by ChatGPT (this is tongue-in-cheek, but it did happen!)-YesChatGPT17/Apr/2023No🔗Patthawikorn Boonrin revealed that he put in a few hypothetical questions... and received the numbers 57, 27, 29, and 99 from the chatbot... the numbers ended up winning him a [lottery] prize...
39
Quantum computingGPT-4 achieves a 'B' grade (73/100) on exam.7374.4NoGPT-411/Apr/2023No🔗The result: GPT-4 scored 73 / 100. (Because of extra credits, the max score on the exam was 120, though the highest score that any student actually achieved was 108.) For comparison, the average among the students was 74.4 (though with a strong selection effect—many students who were struggling had dropped the course by then!). While there’s no formal mapping from final exam scores to letter grades (the latter depending on other stuff as well), GPT-4’s performance would correspond to a solid B.
40
Jurisprudence/
legal rulings
ChatGPT helps a judge with a verdict (India).--ChatGPT29/Mar/2023No🔗‘Armed with [ChatGPT's] legal expertise, Chitkara ultimately rejected the defendant’s bail bid on the grounds that they did act cruelly before the victim died.’
41
Japan: National Medical Licensure ExaminationBing Chat would achieve 78% [above cut-off grade of 70%], ChatGPT would achieve 38%78YesBing Chat9/Mar/2023Yes🔗‘The accuracy of ChatGPT was lower than prior studies using the United States Medical Licensing Examination. The limited amount of Japanese language data may have affected the ability of ChatGPT to correctly answer medical questions in Japanese... Bing has an accuracy level to pass the national medical licensing exam in Japan.’
42
Spanish medical examination (MIR)Bing Chat would achieve 93%, ChatGPT would achieve 70%, both above cut-off grade93YesBing Chat2/Mar/2023No🔗‘I asked 185 questions, excluding the 25 that required images, which I removed. To balance the exam, I added the 10 reserve questions for the challenges. Out of the 185 questions, Bing Chat answered 172 correctly and failed on 13, resulting in a success rate of 93%’
43
Cover of TIME magazineChatGPT made the 27/Feb/2023 cover of TIME magazine.-YesChatGPT27/Feb/2023No🔗Alan: this one is not really an ability, but definitely an achievement!
44
CEOChatGPT appointed to CEO of CS India.--ChatGPT9/Feb/2023No🔗‘As CEO, ChatGPT will be responsible for overseeing the day-to-day operations of CS India and driving the organization's growth and expansion. ChatGPT will use its advanced language processing skills to analyze market trends, identify new impact opportunities, and develop strategies...’
45
Software dev jobChatGPT would be hired as L3 Software Developer at Google: the role pays $183,000/year.-YesChatGPT31/Jan/2023No🔗https://www.pcmag.com/news/chatgpt-passes-google-coding-interview-for-level-3-engineer-with-183k-salary
https://www.cnbc.com/2023/01/31/google-testing-chatgpt-like-chatbot-apprentice-bard-with-employees.html
"ChatGPT gets hired at L3 when interviewed for a coding position"
46
Jurisprudence/
legal rulings
ChatGPT helps a judge with a verdict (Colombia).--ChatGPT31/Jan/2023No🔗English: https://interestingengineering.com/innovation/chatgpt-makes-humane-decision-columbia
Spanish: https://www.bluradio.com/judicial/sentencia-la-tome-yo-chatgpt-respaldo-argumentacion-juez-de-cartagena-uso-inteligencia-artificial-pr30
"On January 31, the first labor court of Cartagena resolved a guardianship action with the help of the famous artificial intelligence known as ChatGPT, arguing that it applied Law 2213 of 2022, which says that in certain cases these virtual tools can be used."
47
PoliticsChatGPT writes several Bills (USA).--ChatGPT26/Jan/2023Yes🔗Regulate ChatGPT: https://malegislature.gov/Bills/193/SD1827
Mental health & ChatGPT: https://malegislature.gov/Bills/193/HD676
48
MBAChatGPT would pass an MBA degree exam at Wharton (UPenn).B/B-YesChatGPT22/Jan/2023Yes🔗"Considering this performance, ChatGPT would have received a B to B- grade on the exam."
49
AccountingGPT-3.5 would pass the US CPA exam.57.6%Yestext-davinci-00311/Jan/2023Yes🔗"the model answers 57.6% of questions correctly"
50
LegalGPT-3.5 would pass the bar in the US.50.3%Yestext-davinci-00329/Dec/2022Yes🔗"GPT-3.5 achieves a headline correct rate of 50.3% on a complete NCBE MBE practice exam"
51
MedicalChatGPT would pass the United States Medical Licensing Exam (USMLE).>60%YesChatGPT20/Dec/2022Yes🔗"ChatGPT performed at >50% accuracy across all examinations, exceeding 60% in most analyses. The USMLE pass threshold, while varying by year, is approximately 60%. Therefore, ChatGPT is now comfortably within the passing range."
52
IQ (fluid/aptitude)ChatGPT outperforms college students on the Raven's Progressive Matrices aptitude test.>98%Yestext-davinci-00319/Dec/2022Yes🔗More info at: https://lifearchitect.ai/ravens/
53
AWS certificateChatGPT would pass the AWS Certified Cloud Practitioner exam.80%YesChatGPT8/Dec/2022No🔗"Final score: 800/1000; a pass is 720"
54
IQ (verbal only)ChatGPT scores IQ=147, 99.9th %ile.>99.9%YesChatGPT6/Dec/2022No🔗"Psychology Today Verbal-Linguistic Intelligence IQ Test, it gets a score of 147!"
55
SAT examChatGPT scores 1020/1600 on SAT exam.52%YesChatGPT2/Dec/2022No🔗"According to collegeboard, a 1020/1600 score is ~52nd percentile."
56
General knowledgeGPT-3 would beat IBM Watson on Jeopardy! questions.100%Yesdavinci20/Sep/2021No🔗Watson scored 88%, GPT-3 scored 100%.
57
IQ (Binet-Simon Scale, verbal only)GPT-3 scores in 99.9th %ile (estimate only).99.9%Yesdavinci11/May/2021No🔗"As of 2021, I expect that it would not be simple to assess the intelligence of an AI using current IQ instrument design... some subtests where the AI would be easily in the top 0.01% of the world population (processing speed, memory), while other subtests may be far lower."
58
General knowledgeGPT-3 outperforms average humans on trivia.73%Yesdavinci12/Mar/2021No🔗"GPT-3 got 73% of 156 trivia questions correct. This compares favorably to the 52% user average."
59
ReasoningGPT-3 would pass the SAT Analogies subsection.65.2%57Yesdavinci28/May/2020Yes🔗"GPT-3 achieves 65.2% in the few-shot setting... average score among college applicants was 57% (random guessing yields 20%)."
60
61
62
See more benchmarks for large language models 2020-2025:
https://lifearchitect.ai/iq-testing-ai/
63
This sheet is owned and maintained by Dr Alan D. Thompson at LifeArchitect.ai. It should be up-to-date, but please see tab date and sheet header.
64
65