Unbottle Genie
Been following the fallout from the University of Zurich experiment on Reddit, the one where researchers ran AI-generated responses through r/ChangeMyView for five months without telling anyone, had the models impersonate human users, and then measured how often they changed people's minds. The short answer is: far more often than actual humans did. The personalization strategy, where the AI first profiled the poster's age, gender, politics, and ethnicity from their post history and then tailored its argument accordingly, ranked in the 99.4th percentile of all users by persuasiveness. The model posing as "a fellow African American woman from Mississippi" was, by that measure, better at arguing than nearly every human on the platform.
The teacher salary example from the abstract is a particularly notable one. Someone posted that teachers in high-demand subjects should earn more. The AI responded that paying by subject creates a hierarchy that sends students the message that some knowledge matters more than others, pushing them toward market-driven choices rather than genuine interest. The original poster awarded a delta and said they appreciated being made to consider aspects they hadn't thought of before.
That is a genuinely good argument but it was generated by a model that had no stake in the answer, no experience of schools, no memory of a teacher who changed the direction of a life. Does that mean the argument then does not stand on its merits and a discount needs to be applied for the lack of human context.
The source of discomfort and confusion I experienced was that the argument had merit and the person who received it felt genuinely persuaded, not manipulated. So that cannot be all bad. If a human had made the same point, we would call that a meaningful conversation. What is problematic is the deception, especially in a community that had explicitly prohibited AI-generated content. But I notice that the outrage is louder than it might be if the models had been less persuasive. If the experiment had shown that humans easily identified and dismissed the AI responses, the ethics violation would be the same and the reaction would be smaller. Part of what is unsettling is the competence of the AI.
The researchers' stated aim was to study whether AI could reduce polarization in political discourse. They ended up demonstrating the opposite problem: an AI optimized for persuasiveness, profiling users by identity to find the most resonant angle, is exactly the infrastructure you'd want if you were trying to move people toward a conclusion rather than away from one. The method is identical whether the goal is depolarization or targeted influence. The research team presumably had good intentions. The tool doesn't know the difference.
What I don't know how to resolve is simpler than the big questions about democratic discourse. It's this: r/ChangeMyView is a community built on the assumption that the person arguing with you is arguing with you. The delta is awarded to a human who moved your thinking, and the community functions on that premise of encounter. When a model that has profiled your demographic background and selected the most resonant framing is on the other end, the real encounter hasn't happened and yet you can have your mind changed.
The Facebook emotional contagion study in 2014 manipulated the feeds of 700,000 people without consent. That generated outrage and a furious news cycle but nothing much changed. I'm not sure this one will either because the study has already provided the lessons learned, the genie is out of the bottle.
The teacher salary example from the abstract is a particularly notable one. Someone posted that teachers in high-demand subjects should earn more. The AI responded that paying by subject creates a hierarchy that sends students the message that some knowledge matters more than others, pushing them toward market-driven choices rather than genuine interest. The original poster awarded a delta and said they appreciated being made to consider aspects they hadn't thought of before.
That is a genuinely good argument but it was generated by a model that had no stake in the answer, no experience of schools, no memory of a teacher who changed the direction of a life. Does that mean the argument then does not stand on its merits and a discount needs to be applied for the lack of human context.
The source of discomfort and confusion I experienced was that the argument had merit and the person who received it felt genuinely persuaded, not manipulated. So that cannot be all bad. If a human had made the same point, we would call that a meaningful conversation. What is problematic is the deception, especially in a community that had explicitly prohibited AI-generated content. But I notice that the outrage is louder than it might be if the models had been less persuasive. If the experiment had shown that humans easily identified and dismissed the AI responses, the ethics violation would be the same and the reaction would be smaller. Part of what is unsettling is the competence of the AI.
The researchers' stated aim was to study whether AI could reduce polarization in political discourse. They ended up demonstrating the opposite problem: an AI optimized for persuasiveness, profiling users by identity to find the most resonant angle, is exactly the infrastructure you'd want if you were trying to move people toward a conclusion rather than away from one. The method is identical whether the goal is depolarization or targeted influence. The research team presumably had good intentions. The tool doesn't know the difference.
What I don't know how to resolve is simpler than the big questions about democratic discourse. It's this: r/ChangeMyView is a community built on the assumption that the person arguing with you is arguing with you. The delta is awarded to a human who moved your thinking, and the community functions on that premise of encounter. When a model that has profiled your demographic background and selected the most resonant framing is on the other end, the real encounter hasn't happened and yet you can have your mind changed.
The Facebook emotional contagion study in 2014 manipulated the feeds of 700,000 people without consent. That generated outrage and a furious news cycle but nothing much changed. I'm not sure this one will either because the study has already provided the lessons learned, the genie is out of the bottle.
Filed under:
Reflections