Skip to content
— CH. 1 · INTRODUCTION —

Google Books

14 min listen · Ch. 1 of 7
7 sections
  • Google Books began not as a product pitch but as a bet placed in a university president's office. When Larry Page visited the University of Michigan and learned that scanning the entire library would take roughly 1,000 years at the current pace, he reportedly told President Mary Sue Coleman that he believed Google could do it in six. That was around 2002, and the ambition behind that claim would go on to generate one of the most consequential and contested archival projects in the history of human culture.

    The goal was audacious: scan every book ever printed. Not some books, not popular books, not books in English. Every book. Google estimated in 2010 that approximately 129,864,880 distinct titles existed in the world, amounting to over 4 billion digital pages and 2 trillion words. By the time Google celebrated the project's fifteenth anniversary, it had scanned more than 40 million titles.

    How did a tech company come to build what might be the largest library ever assembled? What did that effort cost, legally and culturally? And why, after winning a decade-long legal battle, did Google appear to walk away from its own grand vision?

  • When Larry Page and Marissa Mayer first experimented with book scanning in 2002, digitizing a single 300-page book took 40 minutes. The system they eventually built was something else entirely.

    Google established dedicated scanning centers and transported books there by truck. At the heart of each station was a custom-built mechanical cradle that cradled the spine of a book while an array of optical instruments captured the two open pages simultaneously. Two cameras were directed at each page. A LIDAR range finder projected a three-dimensional laser grid across the book's surface to measure the exact curvature of the paper. A human operator turned the pages by hand and triggered the cameras with a foot pedal. At these stations, operators could digitize up to 1,000 pages per hour.

    Many books were scanned using a customized Elphel 323 camera. A patent awarded to Google in 2009 described the system in detail: by constructing a 3D model of each page and then running it through de-warping algorithms that used the LIDAR data, Google could present flat-looking pages without physically flattening them. Traditional flattening methods, such as unbinding books or pressing pages under glass, are damaging at scale. Google's approach protected fragile collections from excessive handling.

    After capture, each page image passed through three levels of processing. First, the de-warping algorithms corrected the page's physical curvature. Next, optical character recognition software transformed the image into searchable text. Finally, a third round of algorithms extracted structural elements: page numbers, footnotes, illustrations, and diagrams. Google also chose to omit color information in favor of higher spatial resolution, since most out-of-copyright books at the time were printed in black and white.

    The speed gains were striking. Scanning operators could eventually process up to 6,000 pages per hour, a rate that made the six-year promise Page made to Coleman at Michigan feel less like fantasy.

  • In December 2004, Google announced partnerships with some of the most storied research libraries in the world: the University of Michigan, Harvard, Stanford, Oxford's Bodleian Library, and the New York Public Library. The plan was to digitize approximately 15 million volumes within a decade.

    The list of partners grew steadily through the years that followed. The University of California system, which manages about 34 million volumes across roughly 100 libraries, joined in August 2006. The Complutense University of Madrid signed on in September 2006, becoming the first Spanish-language institution in the project. The Bavarian State Library followed in March 2007, pledging to make more than a million works available in German, English, French, Italian, Latin, and Spanish.

    Keio University in Japan became Google's first Asian library partner in July 2007, digitizing at least 120,000 public domain books. Mysore University in India announced a partnership to digitize over 800,000 texts, including manuscripts written in Sanskrit and Kannada on both paper and palm leaves dating back to the eighth century.

    By June 2007, the Committee on Institutional Cooperation, later rebranded as the Big Ten Academic Alliance, had committed twelve member libraries to scanning 10 million books over six years. Cornell University Library agreed in August 2007 to digitize up to 500,000 items, both copyrighted and public domain, with Google providing a digital copy of every scanned work back to the university.

    Harvard's holdings alone exceeded 15.8 million volumes. Physical access to those materials was normally restricted to Harvard students, faculty, researchers, and visiting scholars. The project aimed to make discoveries within that collection available to anyone with an internet connection. As of March 2012, the University of Michigan had contributed 5.5 million volumes, and the University of Wisconsin had added about 600,000.

  • Access to scanned books on Google Books is not binary. Google built a layered system of four distinct viewing modes that depends on a book's copyright status and on agreements with publishers.

    Books in the public domain are available in what Google calls "full view": the entire text is readable online and can be downloaded for free. In-print books acquired through the Partner Program can also receive full view status, but only if the publisher explicitly grants permission, which is described as rare. Preview mode applies to in-print books where some permission has been granted; in this case publishers can set the percentage of pages available to read, with a minimum of 20%. Users in preview mode cannot copy, download, or print pages, and a "Copyrighted material" watermark appears at the bottom of each viewable page.

    When Google lacks permission from a copyright owner, users see only a "snippet view": two to three lines of text surrounding the queried search term. If a search term appears many times in the same book, Google caps the display at no more than three snippets to prevent gradual reconstruction of the full text. For certain reference works like dictionaries, Google shows no snippets at all, on the reasoning that even partial text could harm the market for those books. A fourth level, "no preview," applies to books that have not been scanned at all; these appear in search results with only metadata such as title, author, publisher, page count, ISBN, and subject classification.

    For each scanned work, the site generates an overview page that includes publishing details, a high-frequency word map, the table of contents, summaries, and reader reviews. Users with a Google account can export bibliographic citations, write their own reviews, and add books to personal libraries for tagging and sharing. The Ngram Viewer, launched in December 2010, draws on the full collection to graph how the frequency of any word or phrase changed across time periods, a tool that proved useful to historians and linguists.

    Tim Parks, writing in The New York Review of Books in 2014, noted that Google had stopped providing page numbers for many recent publications, a change he attributed to a quiet agreement with publishers designed to push researchers toward buying physical editions.

  • In September 2005, the Authors Guild filed a class-action lawsuit against Google, charging the company with "massive copyright infringement." About a month later, five large publishers and the Association of American Publishers brought their own civil suit. The two cases were eventually consolidated.

    Google's core argument was that its project amounted to fair use: displaying snippets rather than full texts, it said, was the digital equivalent of a card catalog, with the added benefit that every word in every book was indexed. Google also argued it was preserving "orphaned works," meaning books still technically under copyright but whose rights holders could not be located.

    Negotiations produced a proposed settlement in October 2008, but the agreement drew fierce criticism on antitrust, privacy, and class-representation grounds. A federal judge rejected it in March 2011. The publishers settled with Google separately not long after, but the Authors Guild pressed on. In November 2013, US District Judge Denny Chin ruled in Google's favor on fair use grounds. The Authors Guild appealed, and the Second Circuit upheld the ruling in October 2015. The US Supreme Court declined to hear a further appeal in April 2016, letting the lower court's decision stand.

    The project also met resistance outside the United States. In December 2009, a French court ordered Google to stop scanning copyrighted books published in France, awarding French publisher La Martinière and Editions du Seuil 300,000 euros in damages plus 10,000 euros per day until the books were removed from Google's database. The court found that Google had "violated author copyright laws by fully reproducing and making accessible" works without permission. In China, the China Written Works Copyright Society accused Google in late 2009 of scanning 18,000 books by 570 Chinese writers without authorization. Chinese author Mian Mian filed a separate civil lawsuit for $8,900 after Google scanned her novel, Acid Lovers.

    Microsoft's associate general counsel for copyright, Thomas Rubin, accused Google in March 2007 of systematically violating copyright by freely copying any work until told to stop, rather than seeking permission first. Microsoft had funded its own book scanning project that reached 750,000 books and 80 million journal articles before ending it in May 2008.

  • Geoffrey Nunberg, a linguist studying changes in word usage over time, ran a search in Google Books for works published before 1950 that contained the word "internet." The search returned an unlikely 527 results. Woody Allen appeared in 325 books ostensibly published before he was born.

    Metadata errors of this kind run through the collection. A review of the author, title, publisher, and publication year fields for 400 randomly selected Google Books records found that 36% of the sampled books contained metadata errors. Specific examples recorded by scholars include 182 works attributed to Charles Dickens with dates prior to his birth in 1812; an edition of Moby Dick catalogued under "computers"; a biography of Mae West classified under "religion"; ten editions of Whitman's Leaves of Grass simultaneously classified as both "fiction" and "nonfiction"; a title listed as Moby Dick: or the White "Wall"; and, perhaps most dramatically, the metadata for an 1818 mathematical work attached to a completely different book, a 1963 romance novel. Google responded to Nunberg's findings by blaming the bulk of errors on outside contractors.

    Beyond metadata, the scanning process itself introduces physical errors. Pages appear upside down, out of sequence, crumpled, obscured by a scanning operator's thumb, or blurred. Google's own disclosure statement, printed at the end of its digitized books, acknowledges the difficulty plainly: "Smudges on the physical books' pages, fancy fonts, old fonts, torn pages, etc. can all lead to errors in the extracted text." In 2009, Google announced it would use reCAPTCHA to help correct hard-to-read scanned words, though the system could not address errors caused by turned pages or blocked text.

    Language balance drew criticism from European politicians and intellectuals who argued that the overwhelming focus on English-language texts would skew the digital representation of world scholarship. Jean-Noel Jeanneney, the former president of the Bibliotheque nationale de France, was among the most prominent voices raising this concern. German, Russian, French, and Spanish scholarship, critics argued, would be underrepresented in the resulting digital corpus in ways that could shape the direction of future research for decades.

  • After winning the Authors Guild case in 2017, Google might have been expected to accelerate its scanning program. Instead, The Atlantic reported that the company had "all but shut down its scanning operation."

    Wired reported in April 2017 that only a few Google employees remained working on the project, and while new books were still being scanned, the pace had dropped sharply. Librarians at several of Google's partner institutions confirmed that scanning had been slowing since at least 2012. At the University of Wisconsin, the speed had fallen to less than half of what it had been in 2006. Google's own Google Books timeline page, even as of 2017, still listed nothing after 2007, and the Google Books blog had been folded into the Google Search blog back in 2012.

    Google offered no public explanation. Librarians suggested the slowdown might reflect the natural maturation of the project: in the early years, entire stacks of unscanned books could be swept up at once. Later, only the remaining unscanned titles needed attention, making each additional volume incrementally harder to find and process.

    A 2023 study by scholars from the University of California, Berkeley, and Northeastern University's business schools offered an unexpected counterpoint to the mood of retreat. Their research found that Google Books' digitization work had led to increased sales of physical copies of digitized books. The act of making a text searchable and discoverable, it appeared, sent some readers back to the printed page.

    Meanwhile, companion projects had built on what Google started. The HathiTrust Digital Library, launched in October 2008 by the Committee on Institutional Cooperation and the University of California system libraries, was built partly to archive and provide academic access to books that Google and others had scanned. By the time Google's scanning ambitions had visibly dimmed, HathiTrust held about 6 million volumes, more than 1 million of which were in the public domain.

Common questions

What is Google Books and how does it work?

Google Books is a service from Google that searches the full text of books and magazines that Google has scanned, converted to text using optical character recognition, and stored in its digital database. Books are provided either by publishers and authors through the Google Books Partner Program or by libraries through the Library Project. Results appear in both Google Search and the dedicated books.google.com website.

How many books has Google Books scanned?

Google has scanned more than 40 million titles, a figure the company confirmed when it celebrated 15 years of Google Books. Google estimated in 2010 that there were approximately 129,864,880 distinct titles in the world and stated its intention to scan all of them, representing over 4 billion digital pages and 2 trillion words.

Who sued Google over Google Books and what happened?

The Authors Guild filed a class-action lawsuit against Google in September 2005, and five large publishers filed a separate civil suit in October 2005, both citing copyright infringement. The publishers eventually settled separately, while the Authors Guild case continued through multiple appeals. US District Judge Denny Chin ruled in Google's favor in November 2013, the Second Circuit upheld that ruling in October 2015, and the US Supreme Court declined to hear a further appeal in April 2016.

When did Google Books begin and who started it?

Google co-founders Sergey Brin and Larry Page conceived the idea while graduate students at Stanford in 1996. A team at Google officially launched the secret books project in 2002, beginning with scanning experiments that initially took 40 minutes to digitize a single 300-page book. The project was first publicly introduced as Google Print at the Frankfurt Book Fair in October 2004.

What are the different access levels on Google Books?

Google Books has four access levels. Full view allows complete reading and free download of public domain books and some publisher-approved titles. Preview limits viewable pages to a percentage set by publishers, with a minimum of 20%. Snippet view shows two to three lines of text around a search term for books where no permission has been granted, capped at three snippets per term. No preview applies to books that have not been digitized, displaying only metadata such as title, author, and ISBN.

How accurate is Google Books metadata?

A review of 400 randomly selected Google Books records found that 36% contained metadata errors, a rate described as higher than one would expect from a typical library catalog. Documented errors include works attributed to Charles Dickens with dates before his birth in 1812, a Moby Dick edition catalogued under "computers," and the metadata for an 1818 mathematical work incorrectly attached to a 1963 romance novel.

All sources

136 references cited across the entry

  1. 3Read Complete Magazines Online in Google BooksMark O'Neill — 28 January 2009
  2. 5NewsGoogle project promotes public goodKevin Bergquist — University of Michigan — 2006-02-13
  3. 6Is This the Renaissance or the Dark Ages?Andrew K. Pace — American Library Association — January 2006
  4. 815 years of Google Books17 October 2019
  5. 10NewsGoogle Books: A Complex and Controversial ExperimentStephen Heyman — 28 October 2015
  6. 11MagazineWhat Ever Happened to Google Books?11 September 2015
  7. 14Google's Cookie and Hacking Google PrintGreg Duffy — March 2005
  8. 15JournalThe Google Library Project: Both Sides of the StoryJonathan Band — University of Michigan — 2006
  9. 16NewsIn Google Book Settlement, Business Trumps IdealsJuan Carlos Perez — October 28, 2008
  10. 18References, PleaseTim Parks — 13 September 2014
  11. 19Torching the Modern-Day Library of AlexandriaJames Somers — The Atlantic — 20 April 2017
  12. 20Weekly Google Code Roundup for August 10thDion Almaer — 11 August 2007
  13. 22NewsScan This Book!Kevin Kelly — May 14, 2006
  14. 23Patent reveals Google's book-scanning advantageStephen Shankland — 4 May 2009
  15. 24NewsThe Secret Of Google's Book Scanning Machine RevealedMaureen Clements — 30 April 2009
  16. 26Is Google leading an e-book revolution?Laura Miller — 8 December 2010
  17. 34MagazineThe Artful Accidents of Google BooksKenneth Goldsmith — 4 December 2013
  18. 35The trouble with Google BooksLaura Miller — 9 September 2010
  19. 38JournalAn Assessment of Google Books' MetadataRyan James et al. — 2012
  20. 39NewsGoogle's Book Search: A Disaster for ScholarsGeoffrey Nunberg — August 31, 2009
  21. 40BookGoogle and the Myth of Universal Knowledge: A View from EuropeJean-Noël Jeanneney — University of Chicago Press — 2006-10-23
  22. 41NewsFrance Detects a Cultural Threat in GoogleAlan Riding — 2005-04-11
  23. 46Michigan Digitization ProjectUniversity of Michigan
  24. 52Austrian Books OnlineAustrian National Library
  25. 53Google Book Search GrowsAndrew Albanese — 2007-06-15
  26. 64Google to digitise books at Mysore varsityHindustan Times — 20 May 2007
  27. 73NewsA new chapterOctober 30, 2008
  28. 74Authors Guild Sues Google, Citing "Massive Copyright Infringement"Paul Aiken — Authors Guild — 2005-09-20
  29. 75Publishers sue Google over book search projectAlorie Gilbert — CNET News — 2005-10-19
  30. 77Judging Book Search by its coverJen Grant — November 17, 2005
  31. 86Google Book Search Project - MenuBig Ten Academic Alliance
  32. 89Share and enjoyManas Tungare
  33. 92NewsMicrosoft Will Shut Down Book Search ProgramMiguel Helft — May 24, 2008
  34. 93NewsSome Fear Google's Power in Digital BooksNoam Cohen — February 1, 2009
  35. 96NewsGoogle Hopes to Open a Trove of Little-Seen BooksMotoko Rich — January 4, 2009
  36. 100NewsPreparing to Sell E-Books, Google Takes on AmazonMotoko Rich — 2009-06-01
  37. 101NewsFrench court shuts down Google Books projectGaelle Faure — December 19, 2009
  38. 114Google Begins to Scale Back Its Scanning of Books From University LibrariesJennifer Howard — The Chronicle of Higher Education — 9 March 2012
  39. 115MagazineHow Google Book Search Got LostScott Rosenberg — 11 April 2017
  40. 116MagazineGoogle and the Future of BooksRobert Darnton — February 12, 2009
  41. 117NewsAuthors sue Google over book plan21 September 2005
  42. 122NewsFrench publishers toast triumph over GoogleAdam Sage — The Times of London — December 19, 2009
  43. 123NewsGoogle's French Book Scanning Project Halted by CourtHeather Smith — Bloomberg — December 18, 2009
  44. 124NewsFrench publisher sues GoogleJohn Oates — June 7, 2006
  45. 125NewsFine for Google over French booksDecember 18, 2009
  46. 128NewsMicrosoft Attorney Accuses Google Of Copyright ViolationsThomas Claburn — March 6, 2007
  47. 133Home
  48. 137MagazineEurope's Answer to Google Book Search Crashes on Day 1Chris Snyder — November 20, 2008