Timothy Paustian, a biology professor at the University of Wisconsin at Madison, has explored numerous strategies to prevent his students from submitting essays generated by artificial intelligence. Initially, after the introduction of ChatGPT, he experimented with AI detection tools that claimed to identify AI-created text, but found them to be “comically bad.” Subsequently, he incorporated discreet prompts into his assignments designed to elicit responses only from chatbots, such as instructing students to include references to blueberries. Recently, however, he has noticed improvements in some AI detection tools, enabling him to identify 60 AI-composed essays in one assignment from a total of 350 students in his online microbiology class. Nearly all students confessed to using AI, except for one, who opted not to challenge Paustian’s decision to assign a zero for the work.
Now, he finds himself back at square one. Students have adapted to the hidden prompts, evading detection, and his university advises against solely depending on AI detectors. As a result, this year Paustian plans to eliminate writing assignments altogether, expressing his disappointment since he believes writing is an essential learning tool.
Entering the fourth year of the ChatGPT revolution, educational institutions appear no closer to tackling AI-driven cheating than they were at its inception. Chatbots have grown increasingly adept at handling complex tasks, complicating professors' ability to discern machine-generated content. While AI detection technology has seen significant advancements and is gaining traction, many academic circles remain hesitant about its application. Recently, a report from an MIT working group cautioned against over-reliance on these tools, citing concerns about fostering distrust between educators and learners.
Professors find themselves in a challenging position, often skeptical of detection methods yet unsure of alternative strategies. “I think most faculty feel entirely overwhelmed by this,” remarked Marc Watkins, a lecturer at the University of Mississippi focused on AI’s impact in education. “It's chaotic.”
Amid this turbulence, a tool named Pangram has emerged as a leading AI detector, though Turnitin remains the most widely utilized option in academia, offering AI detection as a feature alongside its prominent plagiarism check. Approximately 1,400 colleges and universities in North America have acquired Turnitin’s detection capabilities, which automatically analyzes student submissions. Annie Chechitelli, the company’s chief product officer, indicated that Turnitin identified instances of AI-generated content in nearly half of all student submissions last academic year. While the tool's developers assert a low false positive rate of under 1%, they also acknowledge a false negative rate of approximately 15%. This means there is a chance AI text might be misidentified as human writing, an issue some institutions, like Vanderbilt University, have faced after discontinuing its use due to mistakes.
Pangram, boasting a claimed false positive rate of one in 25,000 from internal assessments, shows potential as a more effective alternative. Research from the University of Chicago last year indicated that Pangram rarely misidentified extensive human-written passages as AI-generated text, though it can err with shorter excerpts. It integrates with course management systems like Canvas, underscoring its growing relevance in educational settings.
Despite Pangram's promising features, it hasn't seen significant adoption in academic environments, which may be attributed to the typically cautious nature of universities. Many institutions, including Wisconsin and Vanderbilt, base their decisions regarding AI detection on research from 2023 that labeled early detectors ineffective, and have been slow to adapt to new findings. Critics have pointed out that previous studies revealed biases in detection tools against non-native English speakers; however, both Turnitin and Pangram contest this claim, asserting their models have undergone improvements.
Consensus among stakeholders— including creators of these detection systems—suggests that AI tools should not be treated as the definitive evidence of academic dishonesty. Chechitelli of Turnitin expressed the importance of using detection results as a starting point for dialogue between educators and students rather than as indisputable proof of cheating. Misguided accusations stemming from these technologies can have severe repercussions for students, with some even taking legal action against universities that penalized them based on AI detection outcomes. Additionally, educators must navigate student privacy laws when submitting essays to these commercial detectors.
A recent paper from researchers at Notre Dame has highlighted another challenge: reliable AI detectors can still be deceived by tools designed to alter text, making it appear less AI-generated. For instance, a student might self-write an essay then employ software like Grammarly to modify it, potentially landing in hot water compared to peers who solely rely on AI for their work.
Conversely, a new perspective is emerging among some educators and administrators who believe students should embrace AI for certain tasks. Indiana University’s Kelley School of Business has posited this view in its “AI playbook,” advocating for assignments that promote thoughtful engagement rather than simply catching AI use. A Harvard dean recently suggested that universities should step away from AI detection and instead support its responsible integration into writing-focused courses.
On a broader scale, the MIT report emphasized that the rise of AI requires more substantial reforms. It warned against “cognitive surrender,” where students might overly depend on machines for thinking. To combat this, schools may need to redesign educational experiences and reassess grading systems to promote creativity and meaningful learning.
While rethinking education may be the ideal approach to AI challenges, it places a considerable burden on professors, especially those in institutions like Ole Miss that may not have the same resources as MIT. Instructors often seek to avoid major academic integrity issues, with Paustian noting that AI misuse hasn’t been substantial in his advanced classes but poses significant challenges in his introductory course for hundreds of students. The underlying issue is less about a lack of solutions and more about the limited time and resources available for educators to implement change effectively.
