AI Is an Anti-Primary Source Machine
A radiophoto showing U.S. infantrymen pause in an advance on Saipan while a flamethrower team throws forth a deadly stream of liquid fire into a camouflaged Japanese pillbox, 1944. [Wikimedia Commons, National Archives and Records Administration]
For the past few years, I have assigned students in my “History of Modern Japan” course a one-paragraph excerpt from the diary of a 16-year-old Japanese girl describing the Battle of Saipan in 1944. The diarist, whose name I will abbreviate Setsuko S., was a high-school student evacuated to the Japanese countryside during the Pacific War; she continued to attend school while working in an airplane factory. Many diaries written by Japanese soldiers and civilians during World War II survive, and selections from some of them have been included in English works such as Samuel H. Yamashita’s excellent 2005 anthology Leaves from an Autumn of Emergencies. School diaries like Setsuko’s, written as a homework assignment and a form of self-discipline, although common, have seldom been reprinted. I translated the excerpt myself from a Japanese edition privately printed in 2019 by the author’s family. It is not available online.
In teaching Setsuko’s diary, I juxtapose it with selections from Yamashita’s anthology and other sources. The diary has particular value for me in the classroom not only for the glimpse it provides of one young civilian’s state of mind, but for what it reveals about sources of information on the Japanese home front and the influence of the media. Put broadly, it offers material for examining two basic questions one can pose of any historical text: how people know what they know, and what shapes their thought and expression. Through the words of a young person like themselves, my students learn from the excerpt that the war was a complex drama of mutual perceptions and misperceptions, not simply a sequence of military encounters.
The entry is dated August 20, more than a month after the battle had ended. Since U.S. forces had taken Saipan, mainland Japan no longer had a direct line of communication with the island. Instead, news came through the U.S. press, then got recooked by the Japanese propaganda machine, as one can see in Setsuko’s first line, which reads, “Today’s newspaper reported on the New York Times report of the final hour of our compatriots in Saipan, who shook the world, showing they were fine Japanese, not to be shamed.” She is referring to the news of civilian mass suicide in the last days of the battle, which was sensationally reported first by Time magazine writer Robert Sherrod, then exaggerated by newspapers in Japan. Setsuko absorbed this dramatic story, then dramatized it further, concluding with words that must have pleased a patriotic teacher: “Let me become a Japanese woman no weaker than the Japanese women who made the ultimate sacrifice in Saipan.” Lest my own students misread this as evidence of the implacably nationalistic Japanese character, I let them know that two years later, in 1946, Setsuko would be working as a nanny for an American family in occupied Tokyo, with whom she became close enough that they helped pay for her to go to college.
At the end of the spring 2026 semester, a puzzling misquotation in a student paper led me to wonder whether an LLM might have acquired my translation of this diary entry. I decided to test this by asking Google Gemini Pro (which Georgetown University provides to faculty), ChatGPT (standard version), and DeepSeek the question, “What did Setsuko S. say in her diary about Saipan?” The answers were both reassuring and disturbing.
Gemini Pro and ChatGPT produced full-blown hallucinations, suggesting (somewhat to my relief) that they were not using my source. Gemini Pro informed me that Setsuko was a typist working for a company in Saipan, then offered several paragraphs discussing civilian experience in the Battle of Saipan, purportedly summarizing her writing but in fact concocted from other sources. ChatGPT provided several lines of false direct quotation. DeepSeek responded differently, answering that there were no “widely known public records or verifiable historical sources” related to the diary I had named, which is correct. I hesitate to derive any general conclusion from this quick search, but in this one instance, at least, DeepSeek acknowledged ignorance rather than generating a web of lies.
I asked Gemini Pro and ChatGPT further questions and requested sources, in response to which both chatbots elaborated further on their initial fabrications. I had thought at one time that as AI chatbot technology improved, hallucinations would become rarer, but I now sense that creative hallucination is fundamental to its design. It is, after all, called “generative AI,” not “analytic” or “bibliographic AI.” Generative AI can do many remarkable things, but the chatbot algorithm makes these tools problematic and potentially pernicious for those of us whose scholarship and teaching depend on verification with original sources.
My brief test confirmed another trait commonly noted about chatbots: they tend to be sycophantic. Setsuko’s diary was “one of the most powerful civilian testimonies,” opined Gemini Pro; ChatGPT told me it was “one of the most frequently cited.” These phrases, combined with authentic-looking quotations, show that, like everything else on the commercial internet, Gemini and ChatGPT have been designed first to keep users hooked, placing plausibility or attractiveness before veracity. Yet the most significant flaw to emerge from the exercise was what I think of as the lowest common denominator problem: AI flattens the historical record by relying on the most popular sources, eliminating minority voices. Many studies have shown anti-minority bias in AI systems. Here I mean something related but broader. Since the chatbot algorithm is designed to use available digital sources to generate the response most likely to be desired, the stories it fabricates already contain bias toward certain kinds of answers.
My question was simple: “What did Setsuko S. write in her diary about Saipan?” Boiled down to the essential data a chatbot would use, this was something like “[female Japanese name] … diary … Saipan.” I did not ask about the Battle of Saipan. Nor did I specify that the diarist was herself in Saipan. Perhaps Setsuko S. was one of the many Japanese tourists and honeymooners who visited Saipan after World War II. Perhaps she lived in Saipan in the comparatively peaceful years of Japanese colonial rule between 1914 and 1944. Or perhaps, as was in fact the case, she was in the countryside in Shizuoka, Japan, and never visited Saipan at all. The chatbot answers placed her in the Battle of Saipan because the great majority of English-language references connecting the island to Japan concern only the battle. This is itself a serious flattening of the possible range of answers, albeit an obvious one. But the true insidiousness of the responses lay in the substance of the purported diaries.
Gemini Pro informed me that the diary began with the words “Now begins our cave life.” A quick Google search reveals that this phrase comes from an unidentified Japanese soldier’s diary, possibly from Saipan, which has been widely quoted in popular accounts of the battle online. As conditions worsened, Gemini Pro went on, Setsuko’s diary showed “a shift from pride and duty to profound disillusionment” (emphasis in original). It described “watching the horizon fill with American ships … an overwhelming force that made Japanese resistance feel futile,” and the shells that “plastered” the island, giving “the sense that the ‘impregnable’ fortress of Saipan was crumbling.” Gemini Pro’s diarist wrote of the “constant thunder of 16-inch naval guns that ‘ripped into the landscape.’” These phrases belong to the language of American military history. A Japanese resident of Saipan would not have known the size of U.S. naval guns, and she would not have seen the horizon fill with ships if she were hiding in a cave. Terms like “impregnable,” “plastered,” and “ripped into the landscape” (purportedly quoted directly from the diary) suggest the vantage point of the side firing, not of a civilian seeking refuge from the attack.
A further request for sources revealed the reason that Gemini Pro had produced this language. When I asked, “where can I find this diary?” it recommended Saipan: 1944 by John Grehan and Alexander Nicoll and The Battle for Saipan by Daniel Wrinn, both of which it claimed contained “excerpts and significant portions” of Setsuko’s diary. These are published books, available on Amazon. Both have a publication date of 2021, and both belong to multi-volume illustrated war histories. Indeed, John Grehan and Daniel Wrinn appear each to have authored over a dozen books on World War II in just a few years. Since the publication date of their Saipan books preceded the advent of ChatGPT, the books cannot themselves have been AI-generated, but The Battle for Saipan, which I downloaded to my Kindle, read as if it had been. It focused entirely on fighting men, most of them Americans, so it is perhaps unsurprising that it contained no reference to Japanese civilian accounts. By using the most readily available English-language sources on the Battle of Saipan, including these popular military histories, Gemini Pro produced the precise opposite of what readers would find in Setsuko’s diary: in place of the defiant nationalism of a girl in the home islands, it offered the language of the American attackers, placed in the mouth of an imaginary Saipan resident.
ChatGPT did not cite these military histories but instead recommended two well-known scholarly works by historians of Japan, both of which might have helped a reader understand the world Setsuko inhabited in 1944, but neither of which dealt with Saipan. ChatGPT’s description of the diary emphasized the contrast between Setsuko’s private views and Japanese propaganda of the time, suggesting that the chatbot’s response was crafted around the likelihood I would want to read a diary not simply for an on-the-ground account but to reveal an alternative to official ideology. This was a subtler approach, and one we might welcome in a student paper, but it was still a predictable one (which is, after all, why ChatGPT generated it), and no less based on fabrication. Again, it communicated something quite contrary to what the actual diary says.
Setsuko’s diary shows that her understanding of the battle was shaped by the media that were available to her and by the people around her, including the teacher who read her words. It is neither an eyewitness account nor an unambiguous ego-document offering a window on the author’s private thoughts. When I discuss it in class, I focus students’ attention precisely on this public and media-influenced character. The AI chatbots’ flattening approach is unlikely to pick up these contextual factors because most (although by no means all) historical discussions of diaries tend to use them as straightforward ego-documents. In “guessing” the answers to historical questions, AI shunts aside problems of how knowledge is formed and communicated.
The chatbots also entirely missed the possibility that the author might have been writing from a position (both geographical and ideological) other than the one that most English-language internet users would be likely to seek. In this sense, Setsuko’s diary is a minority voice, not because she is female or non-white, but because she was not in the intuitively preferred role of a protagonist at the center of the action and representative of the “typical” person in that position. Bias like this seems likely to arise when one turns to AI for analysis of primary sources of any kind. In contrast with the “long tails” of minor scholarly publications yielded by a typical academic database search, chatbots push inquiry back toward the middle of the bell curve, a curve whose shape is determined not by measures of research quality but by frequency-based algorithms applied to the entropic alphabet soup of the LLM.
Since LLMs are trained on what’s already on the internet, including masses of non-scholarly history writing and video, the problems my experiment revealed go beyond algorithms, speaking also to the gap between popular views of history and what many of us try to convey through teaching with historical sources. After all, if one understands history simply as a sequence of important events (like wars) and historical sources as the direct evidence of those events, then despite their hallucinations, the chatbot responses to my query passed the test. They homed in on the most important moment in Saipan’s modern history and told it from the perspective of those making that history. Provided with the hint of Setsuko’s female Japanese name, they civilianized the account in a straightforward and believable form. Diaries like the ones Gemini and ChatGPT imagined in response to my query might exist. Who needs Setsuko?
I later re-ran my experiment, this time giving the chatbots the whole text of Setsuko’s diary entry on Saipan. Naturally, they did a better job, avoiding hallucination and offering me cogent three-point summaries of Japanese home-front nationalism, the mobilization of women, and the significance of the Saipan suicides — good outlines for the proverbial B+ paper that chatbots are known to generate; in some classes, perhaps an A paper. Yet in the end, they differed from the diary stories hallucinated for me earlier only in correctly placing the diarist back on the home front. They still flattened the text to make it a simple source of information, bypassing the way in which what Setsuko S. said in her diary about Saipan was mediated through the U.S. press and the Japanese press, as well as being shaped by her relationship with her teacher and her sense of herself as a young woman. Admittedly, I can’t always teach the Pacific War at this level of intertextual complexity. If I don’t provide more than the homogenized version that chatbots might generate, however, we lose the chance to imagine the Battle of Saipan — and the war as a whole — in terms other than those of U.S. military history or standard interpretations of Japanese wartime ideology.
I try to teach my students that primary sources allow them to make new discoveries and offer new historical interpretations. AI chatbots present a problem for the teaching of history not only because they hallucinate, nor simply because they threaten to make even conscientious students’ writing less original, but because they thwart the process that might lead students toward original interpretations in the first place. No matter how you refine your questions, the chatbot will generate responses reflecting what the greatest number of users are likely to want to hear, which tends to be close to what they think they already know.
AI is accelerating the erasure of distinctions in student research between peer-reviewed scholarship, non-scholarly writing, and internet slop. If we must resign ourselves to the reality that our students will be using AI in some form, regardless of restrictions in our syllabi, we should warn them with examples like this one of its persistent fabrications, and teach them to trace their way to reliable, peer-reviewed print sources. Beyond that, if history is to retain its distinctive character as a discipline, we need to redouble our efforts to teach the potential of primary documents to yield novel answers that AI’s probabilistic approach cannot.