Without an agreed definition, recognising AGI is difficult. Several approaches have been proposed.
The Turing Test
Alan Turing proposed judging whether a machine's conversation is indistinguishable from a person's. Modern chatbots can often pass casual versions, which suggests the test measures conversational imitation more than general intelligence.
Benchmark Suites
Collections of hard tests in maths, science, coding and reasoning. Models have rapidly saturated many benchmarks, prompting ever harder ones. Benchmark success doesn't always translate to real-world reliability.
Novel Problem Solving
Tests designed so that memorisation doesn't help, measuring the ability to learn new skills from few examples.
Economic Measures
Whether AI can perform a large share of real jobs or economically valuable tasks end to end.
Autonomy Measures
How long and complex a task an AI can complete independently — hours, days, weeks of human-equivalent work.
Levels Frameworks
Some researchers propose graded levels of generality and performance rather than a single line, similar to levels used for self-driving cars.
The Takeaway
No single test settles it. Look at multiple capabilities, reliability and real-world performance together.