The Paradox of Digitization: Why Is AI Destroying Books After Scanning Them? / 디지털화의 역설: AI는 왜 책을 스캔한 뒤 원본을 파괴하는가
When Digitization Began as a Way to Preserve Knowledge / 지식을 보존하려고 시작한 디지털화
The Paradox of Digitization: Why Is AI Destroying Books After Scanning Them?
When Digitization Began as a Way to Preserve Knowledge
For a long time, we have believed that digitization is a technology of preservation.
Paper books deteriorate, burn, become water-damaged, and disappear with time. Digital files, by contrast, can be copied indefinitely and accessed from anywhere in the world. Digitizing old and rare books therefore seemed to be one of the safest ways to transmit humanity’s knowledge into the future.
Yet a recent report from the United Kingdom raises a troubling question.
According to the report, some AI companies have been purchasing large quantities of used and rare books around the world, digitizing them for AI training, and then discarding or shredding the original physical books.
The image accompanying the report is particularly striking: an old book is being fed into a machine and reduced to scraps of paper, while an AI symbol appears beside the digitization process.
A technology once used to preserve human knowledge may now be contributing to the destruction of the physical originals of that knowledge.
This is not simply a question of what happens to a few old books.
It is a civilizational question about what we believe knowledge is in the age of AI.
What Is a Book to AI?
From the perspective of an AI company, a book is a source of data.
A book written hundreds of years ago may be a cultural artifact to a human being. To an AI system, however, it may contain hundreds of thousands of words, concepts, expressions, relationships, and patterns that can contribute to model training.
AI also “reads” books differently from human beings.
A human reader can experience the physical texture of a book, examine its cover, notice the quality of its paper, and interpret handwritten notes and other traces left by previous readers as part of its cultural meaning.
For AI, however, what matters primarily is the digital representation of the text.
From this perspective, buying a book, scanning it, and discarding the original may appear economically rational.
But this is precisely where efficiency collides with cultural value.
AI may need data.
Humanity needs more than data.
Humanity needs memory.
A Book Is More Than Its Text
Consider an old book.
It contains more than the words written by its author.
It also contains the printing technology of its era, the characteristics of its paper, its binding techniques, and traces left by the people who owned and read it.
Even a small handwritten note in the margin can become valuable evidence for a historian.
A sentence written in the margin may reveal the reading culture of a particular period. A printing error in a particular edition may help scholars reconstruct the history of publishing and distribution.
Therefore, a digital copy and an original book are not necessarily the same thing.
Digitization can preserve the content of a book, but it does not necessarily preserve every aspect of the historical context embodied in the physical object.
For a rare book, the original itself may be historical evidence.
If the original is destroyed after scanning, we may preserve its content while permanently losing an important physical cultural artifact.
AI Advancement Must Not Come at the Cost of Cultural Memory
This leads to a more fundamental question:
How much of humanity’s cultural heritage should be consumed for the advancement of AI?
AI is growing on the enormous body of knowledge accumulated by humanity. Literature, history, philosophy, science, art, and countless individual creations can contribute to the intellectual foundation of AI systems.
If AI uses this knowledge, should there not also be a corresponding responsibility concerning how that knowledge is acquired and preserved?
The destruction of a rare or unique original after digitization is not merely a matter of corporate efficiency.
It may deprive future generations of the opportunity to make their own judgments about the value of that original.
Something we consider unimportant today may become essential evidence for researchers fifty years from now.
Throughout history, enormous amounts of human knowledge have disappeared through war, fire, and natural disasters. The loss of countless manuscripts—including the legendary loss associated with the Library of Alexandria—has left gaps in humanity’s intellectual record that can never be completely filled.
If, in the age of AI, we begin to say, “We digitized it, so we no longer need the original,” we may unintentionally create a new form of knowledge loss.
Digitization Should Be the Beginning of Preservation, Not the End
Digitization itself is not the problem.
On the contrary, digitization is one of the most powerful technologies humanity has ever developed for preserving and sharing knowledge.
The real question is what we do after digitization.
The most sensible principle is simple:
Whenever possible, digitize the material—and preserve the original as well.
Rare books, unique editions, and historically significant documents should ideally be preserved in both digital and physical forms.
If an AI company purchases large numbers of books, it could cooperate with libraries, universities, museums, archives, and other preservation institutions so that the physical originals can be donated or placed in long-term custody.
AI training requires digital data. It does not necessarily require the destruction of paper.
If alternatives are technologically feasible, destroying originals merely for economic convenience is not fundamentally a technological problem.
It is a problem of values.
The Need for a New Digital Cultural Heritage Ethics
Until now, we have tended to regard cultural heritage preservation primarily as a responsibility of museums, libraries, archives, and governments.
Generative AI changes this situation.
AI companies are becoming some of the world’s largest collectors and processors of knowledge.
They increasingly process books, web pages, images, music, video, academic papers, software code, and countless other forms of human intellectual production.
This means AI companies may eventually need to accept new responsibilities.
How should cultural materials acquired for AI training be preserved?
Who should be responsible for preserving digitized originals?
What standards should govern the disposal of rare materials?
How can we guarantee future generations the right to access important originals?
These questions may become an important part of AI governance.
AI ethics is not merely about preventing AI systems from harming human beings.
It is also about how AI should treat human civilization.
From an AI That Reads to an AI That Remembers
We should not think only about teaching AI to read books.
We should also think about how AI can help preserve humanity’s memory.
Future AI systems could move beyond simply learning the content of books. They could preserve information about their sources, editions, historical backgrounds, ownership histories, and physical characteristics.
Imagine digitizing a rare book and preserving not only its text but also high-resolution images, edition information, publication date, publisher, provenance, physical condition, and other metadata.
This would be more than a database.
It could become a digital memory institution for human knowledge.
Generative AI and digital humanities can play an important role here. AI can help translate ancient documents, search enormous collections, connect seemingly unrelated texts, and generate new research questions.
But the existence of the original materials should not disappear in the process.
AI should be a technology that expands memory—not a machine that consumes memory.
An Important Lesson for Korea
This issue is particularly relevant to Korea.
Korea possesses centuries of accumulated cultural heritage, including classical books, manuscripts, woodblocks, genealogical records, literary collections, Buddhist texts, and personal documents.
Many of these materials are not merely containers of information. Their particular editions and physical forms can themselves be objects of historical research.
Consider, for example, the collected works of a Joseon-era scholar.
Digitizing such a collection does not make the original irrelevant. In many respects, digitization can make the original even more valuable.
The digital version allows researchers around the world to access and study the work. But the original remains physical evidence that the document existed within a particular historical context.
Korea should therefore consider developing, in the age of AI, national principles for cultural heritage preservation in AI training.
When rare historical documents are digitized for AI-related purposes, preservation of the physical original should be a fundamental principle.
The Most Important Question We Should Ask AI
Perhaps the most important question of the AI era is not simply:
“What can AI learn faster?”
It may instead be:
“What must we never allow ourselves to lose?”
AI can read far faster than humans.
It can process vastly more data and connect information at extraordinary speed.
But no matter how powerful AI becomes, it cannot recreate a unique original once it has been destroyed.
A digital copy may preserve the text, but it does not necessarily restore the historical existence and material authenticity of the original.
The true wisdom of the AI era therefore lies in combining the ability to collect more knowledge with the responsibility to preserve more knowledge.
Technology Should Exist to Preserve Memory
It may take only minutes or hours to scan a book.
But it may have taken decades or centuries for that book to be written, printed, read, preserved, and passed down to us.
We should respect that long journey.
As AI enters an era in which it learns from humanity’s accumulated knowledge, our most important responsibility is not simply to provide AI with more data.
It is to decide what must be preserved.
In the digital age, we can move a book from paper into data.
But the fact that its content has become digital does not mean that the value of the original has also been transferred into the data.
Therefore, let AI read our books—but do not let AI kill the books it reads.
The true measure of technological progress is not how much we can digitize.
It is whether we can wisely determine what should be digitized, what should be preserved, and what must be passed on to future generations.
That may be one of the most important forms of wisdom that humanity—not AI—must learn first in the age of artificial intelligence. +++
{Solti}
August 8, 2026
디지털화의 역설: AI는 왜 책을 스캔한 뒤 원본을 파괴하는가
지식을 보존하려고 시작한 디지털화
우리는 오랫동안 디지털화를 보존의 기술이라고 믿어 왔다. 종이책은 낡고, 불에 타고, 물에 젖고, 세월 속에서 사라진다. 반면 디지털 파일은 복제할 수 있고 세계 어디에서나 접근할 수 있다. 그래서 오래된 책과 희귀본을 디지털로 옮기는 일은 인류의 지식을 미래로 전달하는 가장 안전한 방법처럼 보였다.
그런데 최근 영국에서 전해진 한 소식은 이 믿음에 묵직한 질문을 던진다.
보도에 따르면 일부 AI 기업들이 전 세계의 중고·희귀 서적을 대량으로 구입한 뒤, AI 모델 학습을 위해 책을 디지털화하고 원본을 폐기하거나 파쇄하는 사례가 나타나고 있다. 기사 속 사진은 오래된 책이 거대한 기계에 들어가 종이 조각으로 변하고, 그 옆의 화면에는 ‘AI’가 표시되는 모습을 상징적으로 보여준다.
한때 인류의 지식을 보존하기 위해 사용했던 디지털화 기술이 이제는 오히려 지식의 물리적 원본을 없애는 도구가 될 수 있다는 것이다.
이것은 단순히 책 몇 권이 사라지는 문제가 아니다.
우리가 AI 시대에 지식을 무엇이라고 생각하고 있는가에 관한 문명사적 질문이다.
AI에게 책은 무엇인가
AI 기업의 입장에서 책은 데이터의 원천이다.
수백 년 전에 쓰인 책 한 권은 인간에게는 하나의 문화유산이지만, AI 시스템에게는 수십만 개의 문장, 개념, 표현, 관계와 지식 패턴을 포함한 거대한 학습 데이터일 수 있다.
AI가 책을 읽는 방식은 인간과 다르다. 인간은 책의 물성을 느끼고, 표지를 보고, 종이의 질감을 경험하며, 책에 남은 흔적과 소유자의 메모까지 문화적 의미로 받아들일 수 있다.
그러나 AI에게 중요한 것은 텍스트의 디지털 표현이다.
이러한 관점에서는 책을 구입하고 스캔한 뒤 원본을 버리는 것이 경제적으로 합리적으로 보일 수도 있다.
하지만 바로 여기에서 효율성과 문화적 가치 사이의 충돌이 발생한다.
AI가 필요로 하는 것은 데이터일 수 있지만, 인간에게 필요한 것은 단순한 데이터만이 아니다.
인간에게 책은 기억이다.
책은 텍스트 이상의 것이다
오래된 책 한 권을 생각해 보자.
그 안에는 글쓴이가 남긴 문장만 있는 것이 아니다. 책이 만들어진 시대의 인쇄 기술이 있고, 종이의 재질이 있고, 제본 방식이 있고, 책을 소유했던 사람들의 흔적이 있다.
누군가 연필로 남긴 작은 메모 하나도 역사학자에게는 중요한 자료가 될 수 있다.
책의 여백에 적힌 한 문장은 한 시대의 독서문화를 보여줄 수 있고, 특정 판본의 오탈자는 당시의 출판 기술과 유통 과정을 추적하게 해줄 수도 있다.
따라서 디지털 복제본과 원본은 동일하지 않다.
디지털 파일은 책의 내용을 보존할 수 있지만, 원본의 모든 역사적 맥락을 반드시 보존하는 것은 아니다.
특히 희귀서적이라면 원본 자체가 역사적 증거다.
그것을 스캔한 뒤 파쇄한다면 우리는 책의 ‘내용’은 얻었을지 모르지만, 책이라는 물리적 문화유산 하나를 영원히 잃게 된다.
AI의 발전이 지식의 소멸을 가져와서는 안 된다
여기에서 더 중요한 질문이 나온다.
AI 발전을 위해 인간의 문화유산을 어디까지 소비해도 되는가?
AI는 인류가 축적한 방대한 지식 위에서 성장하고 있다. 문학, 역사, 철학, 과학, 예술 그리고 수많은 개인의 창작물이 AI의 지적 능력을 형성하는 토대가 된다.
그렇다면 AI가 그 지식을 활용하는 방식에도 일정한 책임이 따라야 하지 않을까?
특히 희귀하거나 유일한 원본을 확보한 뒤 디지털화하고 폐기하는 행위는 단순한 기업의 효율성 문제가 아니다.
그것은 미래 세대가 선택할 권리를 빼앗을 수도 있다.
오늘 우리가 중요하지 않다고 생각하는 원본이 50년 뒤에는 새로운 연구의 핵심 자료가 될 수도 있기 때문이다.
인류 역사를 돌아보면 수많은 문명의 기록이 전쟁과 화재와 자연재해로 사라졌다. 알렉산드리아 도서관의 상실을 비롯해 수많은 문헌의 소멸은 오늘날까지 인류의 지적 세계에 빈 공간으로 남아 있다.
그런데 AI 시대에 우리가 스스로 ’디지털화했으니 원본은 필요 없다’고 생각한다면, 우리는 새로운 형태의 지식 소실을 만들어낼 수도 있다.
디지털 복제는 보존의 끝이 아니라 시작이어야 한다
디지털화 자체가 문제는 아니다.
오히려 디지털화는 인류가 지식을 보존하고 공유하는 가장 강력한 기술 가운데 하나다.
문제는 디지털화 이후의 선택이다.
가장 바람직한 원칙은 간단하다.
가능하면 디지털화하고, 동시에 원본도 보존하라.
특히 희귀본과 유일본, 역사적 가치가 높은 자료는 디지털 데이터와 물리적 원본을 함께 보존해야 한다.
AI 기업이 책을 대량으로 구입한다면 도서관이나 대학, 기록기관, 박물관과 협력하여 원본을 공공 또는 전문 보존기관에 기증하거나 장기 보관하는 방법도 가능하다.
AI 학습에 필요한 것은 디지털 데이터이지 반드시 종이의 파괴가 아니다.
기술적으로 가능한 선택지가 있는데도 경제적 편의 때문에 원본을 없앤다면, 그것은 기술의 문제가 아니라 가치의 문제다.
AI 시대에는 새로운 ‘디지털 문화유산 윤리’가 필요하다
지금까지 우리는 문화유산 보호를 주로 박물관과 도서관의 문제로 생각했다.
하지만 생성형 AI의 등장으로 상황이 달라졌다.
AI 기업은 이제 세계에서 가장 거대한 지식 수집자가 되고 있다.
그들은 책뿐 아니라 웹페이지, 이미지, 음악, 동영상, 논문, 코드 등 인간이 만들어낸 수많은 지적 자산을 처리한다.
따라서 앞으로는 AI 기업에도 새로운 책임이 요구될 수 있다.
AI 학습을 위해 취득한 문화유산을 어떻게 보존할 것인가?
디지털화한 원본을 누가 보관할 것인가?
희귀 자료를 폐기할 경우 어떤 기준과 절차가 필요한가?
미래 세대가 원본에 접근할 권리를 어떻게 보장할 것인가?
이런 질문들은 앞으로 AI 거버넌스의 중요한 부분이 될 것이다.
AI 윤리는 단순히 AI가 인간에게 해를 끼치지 않도록 하는 문제가 아니다.
AI가 인간 문명을 어떻게 다루어야 하는가에 관한 문제이기도 하다.
‘읽는 AI’에서 ‘기억하는 AI’로
우리는 AI에게 책을 읽게 하는 것만 생각해서는 안 된다.
AI가 인류의 기억을 보존하는 방법도 함께 생각해야 한다.
미래의 AI 시스템은 단순히 책의 내용을 학습하는 것을 넘어 책의 출처, 판본, 역사적 배경, 소유 이력, 물리적 특성까지 연결해서 보존하는 방향으로 발전할 수 있다.
예를 들어 하나의 희귀본을 디지털화한다면 텍스트만 저장하는 것이 아니라 고해상도 이미지, 판본 정보, 제작 연도, 출판사, 저자, 소유 이력, 보존 상태 등을 함께 기록할 수 있다.
이것은 단순한 데이터베이스가 아니다.
그것은 인류 지식의 디지털 기억기관이 될 수 있다.
여기에서 생성형 AI와 디지털 인문학은 새로운 역할을 할 수 있다. AI는 오래된 문헌을 번역하고, 검색하고, 서로 연결하고, 새로운 연구 질문을 발견하는 데 활용될 수 있다. 그러나 그 과정에서 AI가 학습한 원본의 존재 자체가 사라져서는 안 된다.
AI는 기억을 소비하는 기계가 아니라 기억을 확장하는 기술이어야 한다.
한국에도 중요한 교훈
이 문제는 한국과도 무관하지 않다.
한국에는 수백 년 동안 축적된 고문헌과 고서, 목판, 족보, 문집, 사찰 기록, 개인 문서들이 존재한다.
특히 조선시대 문집처럼 하나의 판본 자체가 역사적 연구의 대상이 되는 자료도 많다.
예를 들어 한 문인의 문집을 디지털화한다고 해서 원본의 가치가 사라지는 것은 아니다. 오히려 디지털화가 이루어질수록 원본의 중요성은 더욱 커질 수 있다.
디지털 기술은 세계 어디에서나 그 문헌을 읽게 해주지만, 원본은 그 문헌이 실제로 존재했던 역사적 증거이기 때문이다.
따라서 한국도 AI 시대를 맞아 국가도서관, 대학도서관, 박물관, 연구기관, AI 기업이 협력하여 ’AI 시대 문화유산 보존 원칙’을 마련할 필요가 있다.
특히 희귀 문헌의 AI 학습을 위해 디지털화하는 경우에는 원본 보존을 기본 원칙으로 삼는 것이 바람직하다.
인류가 AI에게 물어야 할 가장 중요한 질문
AI 시대의 가장 중요한 질문은 어쩌면 이것인지도 모른다.
“무엇을 더 빨리 학습할 것인가?”
가 아닐 수 있다.
오히려,
“무엇을 절대로 잃어버리지 말아야 하는가?”
가 더 중요한 질문일 수 있다.
AI는 인간보다 훨씬 빠르게 읽고, 훨씬 많은 데이터를 기억하고, 훨씬 빠르게 지식을 연결할 수 있다.
그러나 AI가 아무리 뛰어나더라도 한 번 파괴된 유일한 원본을 되살릴 수는 없다.
디지털 복제본이 남아 있다고 해서 원본의 역사적 존재까지 복원되는 것도 아니다.
그래서 AI 시대의 진정한 지혜는 더 많이 수집하는 능력과 더 많이 보존하는 책임을 동시에 갖는 것이다.
기술은 기억을 위해 존재해야 한다
책 한 권을 스캔하는 데는 몇 분 또는 몇 시간밖에 걸리지 않을 수 있다.
그러나 한 권의 책이 만들어지고, 읽히고, 보존되어 오늘 우리에게 도착하기까지는 수십 년, 수백 년이 걸릴 수 있다.
우리는 그 긴 시간을 존중해야 한다.
AI가 인간의 지식을 배우는 시대에 인간이 해야 할 가장 중요한 일은 AI에게 더 많은 데이터를 공급하는 것만이 아니다.
무엇을 보존해야 하는지 결정하는 것이다.
디지털 시대에 우리는 책을 종이에서 데이터로 옮길 수 있다.
그러나 데이터가 되었다고 해서 원본의 가치까지 데이터 속으로 옮겨지는 것은 아니다.
따라서 AI에게 책을 읽게 하되 책을 죽이지 말아야 한다.
AI가 인류의 기억을 학습하는 시대일수록, 인류는 자신의 기억을 더욱 소중히 보존해야 한다.
기술의 진정한 진보는 더 많은 것을 디지털화하는 데 있지 않다.
무엇을 디지털화하고, 무엇을 보존하며, 무엇을 미래 세대에게 반드시 남겨야 하는지를 현명하게 선택하는 데 있다.
그것이 AI 시대에 인간이 AI보다 먼저 배워야 할 가장 중요한 지혜다. +++
2026년 8월 8일
{솔티}
Prof. Dr. Young Choi (Editor in Chief) — Regent University
Young B. Choi is a Professor in the Department of Engineering & Computer Science at Regent University. He published 38 books with ‘Selected Readings in Cybersecurity’ (2018) (over 800 copies archived globally at university/college libraries around the world) and ‘Cybersecurity Applications and Artificial Intelligence’ (2023) available in seven major world languages. He proposed the world’s first global and universal telecommunications “Service Order Handling (SOH)” Model (T-SOH Model) (1995) with Dr. Adrian Tang. With this innovative research work, he received the IEEE NOMS ’96 Best Paper Award and became the first recipient of the Outstanding Contribution Award of the TeleManagement Forum in 1998. His research areas include Natural Language Processing-focused AI, AI-applied cybersecurity, network and telecom service management, and Korean studies on Gani Choi Rip’s Jeonggwan (靜觀: Quiet Contemplation) philosophy and Shilhak ( 實學: Practical Learning).
© K-GSP (K-Global Scholars and Professionals) Forum. All rights reserved. August 2026. Content published in the K-GSP Forum may not be reproduced, distributed, or transmitted in any form without prior written permission from the K-GSP Forum, except for brief quotations with full attribution.



