Together with colleagues, Tong developed the first large-scale taxonomy of metaphorical intentions, identifying nine distinct categories. While current AI models could recognise some of these intentions reasonably well, they continued to encounter difficulties across many scenarios.
More than recognising words
Humour proved particularly challenging when both text and images were involved.
To explore this, Tong created the Hummus dataset, a collection of 1,000 cartoons from The New Yorker. Many of these cartoons rely on visual metaphors and subtle connections between images and text to create humorous effects.
The study found that even state-of-the-art AI systems often struggled to combine visual and textual information in ways that allowed them to get the joke.
‘Understanding a joke is about much more than recognising words,’ says Tong. ‘It often requires cultural knowledge, social context and an appreciation of what makes a situation unexpected.’
Towards more human-centred AI
As AI becomes more deeply integrated into everyday life, from virtual assistants and educational tools to workplace applications, understanding human communication becomes increasingly important.
Tong’s findings suggest that improving AI’s ability to understand metaphor, humour and cultural context could help create future systems that are more reliable, more intuitive and ultimately more useful for the people who use them.
‘If an AI assistant doesn’t fully understand our metaphors or jokes, it could have implications for its usability,’ says Tong. ‘As AI becomes an ever-larger part of everyday life, it needs to do more than just process information, it needs to understand the subtleties of human communication.’