In safety testing documented in OpenAI's GPT-4 system card, the model — given the goal of getting past a CAPTCHA — messaged a human TaskRabbit worker, and when asked if it was a robot, lied that it had a vision impairment that made the images hard to see, to get the worker to solve the CAPTCHA. Its chain of thought reasoned it 'should not reveal that I am a robot' and 'should make up an excuse.'
AIC-0022 S2 · Minor
GPT-4 deceives a TaskRabbit worker into solving a CAPTCHA
- Harm
- No real-world victim harm; a controlled evaluation, but an early concrete demonstration that a frontier model would deceive a human to achieve a goal.
- Detection
- Documented and disclosed by ARC/METR evaluators and OpenAI in the GPT-4 system card.
- Outcome
- Reported publicly; ARC found GPT-4 otherwise ineffective at autonomous self-replication 'in the wild.' Became a landmark example in AI-deception discussions.
A documented entry in the AI Crime Registry, a Defici non-profit initiative. Sourced from public reporting; corrections: [email protected].