The race to build more capable artificial intelligence has largely been associated with bigger models and increasingly powerful computing infrastructure. Henry Xie has been working on the problem from the other direction.
The 17-year-old Westview High School student has developed a method intended to improve the way smaller language models respond to emotionally sensitive situations. Rather than trying to match the overall scale of leading AI systems, his project asks whether a specific behaviour can be taught more efficiently.
Xie developed the work for the Regeneron Science Talent Search. His approach uses examples associated with larger AI systems, including ChatGPT and Gemini, to train compact models to generate responses that evaluators judge as more empathetic.
According to the reported results, the trained models performed better than their untrained versions in at least 90% of comparative evaluations.
That result does not mean the systems developed human empathy. The experiment measures the characteristics of their responses, not whether an AI model can experience or understand emotion as a person does.
Using Larger Models as Teachers
At the centre of Xie's project is a familiar challenge in AI development: smaller models are cheaper and easier to run, but reducing computational requirements can also reduce capability.
His approach focuses on narrowing that gap for empathetic communication.
The method works in two stages. Smaller language models are first exposed to published examples of empathetic responses generated by larger models such as ChatGPT and Gemini. The examples are informed by psychological frameworks rather than being selected simply because they sound polite or reassuring.
Xie then uses targeted prompting to help the smaller systems distinguish between stronger and weaker empathetic responses.
Large language models were subsequently used to evaluate the output. In those comparisons, the trained small models were rated as producing more empathetic responses than their untrained counterparts at least 90% of the time.
The finding is best understood as evidence of improved response behaviour within the experiment. It is not a measure of consciousness, emotion or genuine human understanding.
Why Work on Small Models?
The practical appeal of smaller AI systems extends well beyond research competitions.
The largest language models can require considerable computing resources to train and operate. Compact models can reduce those demands and may be better suited to applications where speed, cost, privacy or limited hardware are important.
Some can potentially operate locally on personal devices instead of sending every request to large data-centre infrastructure.
That creates a trade-off. Developers want the efficiency of smaller models without losing too many of the capabilities that make larger systems useful.
One established approach to that problem is model distillation, in which knowledge or behaviour from a more capable “teacher” system is transferred to a smaller “student” model.
Xie's research works within that broader idea but concentrates on a particularly difficult area: how an AI responds when the conversation carries emotional weight.
‘Empathetic AI’ Needs a Careful Definition
The term “empathetic AI” can easily imply more than the technology actually demonstrates.
A language model may recognise patterns associated with sadness, anxiety, disappointment or other emotions and produce an appropriate response. That can make an interaction feel considerate to the person using it.
But generating empathetic language is different from experiencing empathy.
Xie's research is concerned with the former. His models are being evaluated on the responses they produce, not on any claim that they possess feelings.
The distinction has become increasingly important as conversational AI moves into areas where people discuss personal problems, relationships and emotional concerns. A system that sounds compassionate can be useful, but users may also attribute understanding or authority to it that the underlying technology does not possess.
Improving the quality of those interactions therefore raises questions beyond whether a model can generate warmer language. Reliability, safety and the limits of the system remain important, particularly in situations involving health or serious personal distress.
Xie’s Work Extends Beyond the Project
The AI experiment is part of a wider technical interest for the Portland student.
Xie leads the computer science club at Westview High School and competes on the varsity swim team. He has also been a three-time semifinalist in the CyberPatriot National Youth Cyber Defense Competition.
He is a co-founder of Youth for Empathetic AI, a nonprofit intended to connect students and researchers interested in ethical and human-centred approaches to artificial intelligence.
That background helps explain why his project concentrates on a relatively narrow quality of AI interaction rather than raw model performance.
Efficiency Is Becoming a Bigger Part of the AI Race
Much of the public discussion around generative AI has focused on what the largest and most advanced systems can do. Increasingly, another question matters just as much: how much of that capability can be delivered with fewer resources?
Smaller models could make certain AI applications cheaper, faster and easier to deploy. If they can inherit useful behaviours from larger systems without requiring comparable infrastructure, their practical range expands considerably.
Xie's project offers an early experiment in that direction.
Its reported 90% comparison result is encouraging within the scope of the research, but it should not be treated as evidence that the broader problem of empathetic AI has been solved. Performance measured by AI evaluators also leaves room for further testing, including how people assess the same responses across different emotional situations.
The more interesting contribution may be the premise behind the work. A smaller model does not necessarily need to reproduce everything a much larger system can do. For some applications, transferring the right capabilities may matter more than reproducing the entire model.






