{"id":90757,"topic":"ai","source":"The Guardian","title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","url_hash":"d7a0ae8df66791ce7925bfcf7bc57660560aed05","author":"","summary":"<a href=\"https://news.google.com/rss/articles/CBMipAFBVV95cUxOZDBLbUhWbUtyVUg2bWZsNFF4NDdyRWVydjNwX3ctQldEbkN2cW4ybi1hU0tMdy1jMzlVSGtXQVVHcnEwVmxVVFh2bHk0UXBtMWxqRE5Bc3doanJTcWY5TDBCTXdpZzFJWDk1QUc0UjBTTU5ZWDU2c3NFZW9BRkZhWEd1U0pFUDNFRXpOT3ZOazQ4dHNZcTY5UTVOM3ByWUtHUnBBdg?oc=5\" target=\"_blank\">Oxford lets OpenAI train its AI models on Bodleian library</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">The Guardian</font>","content":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data.\nThe Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.\nOxford announced a partnership with the company in March 2025, using OpenAI software to digitise texts from the university’s world-famous library, which it said would make the content more widely available for students and researchers.\nHowever, the announcement did not state the material would be used for training OpenAI’s models, which are trained to recognise patterns in words – and thus “learn” to write complete sentences and perform other cognitive tasks – by being fed vast amounts of data.\nAn OpenAI spokesperson said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”.\n“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” they added.\nMeeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI and the effect on the university’s environmental commitments of striking a deal involving an energy-intensive technology.\nBooksellers have also reported a spate of orders for obscure titles such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners have speculated that because the titles are unlikely to exist online in digitised form, they represent fresh data that can be consumed by the next generation of AI models.\nScraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections.\nOpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project.\nBy June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other texts scanned include a rare collection of 10,000 16th-century “broadside ballads” containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitising of 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.\nThe OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.\nA spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.\nThey rejected the suggestion that the machine-learning element had been hidden from the public and students. Digitisation was the university’s primary interest, said the spokesperson, but staff had been open that the project would also contribute training data.\nThe Bodleian’s collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning. Anthropic, OpenAI’s close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books.\nA tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned.\nThe Oxford spokesperson said: “The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.\n“The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with outputs from other digitisation partnerships.\n“The project will in fact allow the Bodleian to make the material more accessible to a wider number for people, who might otherwise have found it difficult to access.”","image_url":"https://i.guim.co.uk/img/media/f013338350bc3c6214a44769307b23d3e615a327/333_0_3333_2667/master/3333.jpg?width=1200&height=630&quality=85&auto=format&fit=crop&precrop=40:21,offset-x50,offset-y0&overlay-align=bottom%2Cleft&overlay-width=100p&overlay-base64=L2ltZy9zdGF0aWMvb3ZlcmxheXMvdGctZGVmYXVsdC5wbmc&enable=upscale&s=b74d7555c5ff03c9f1d0748fdc68055e","lang":"en","published_at":"2026-09-26T11:12:00+00:00","fetched_at":"2026-09-26T12:15:05+00:00","status":"read","starred":0,"extract_state":"ok","summary_auto":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.","cluster_id":null,"extract_retries":0,"extract_error":null,"contract_version":"news_item.v1","format_contract_version":"news_item_formats.v1","dedup_url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4764 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4764,"summary_length":323,"usable_text_length":4764,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4764,"summary_length":323}},"news_item":{"id":90757,"canonical_url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","source_url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","source_name":"The Guardian","author":null,"published_at":"2026-09-26T11:12:00+00:00","locale":"en","topic":"ai","tags":[],"rss_summary":"<a href=\"https://news.google.com/rss/articles/CBMipAFBVV95cUxOZDBLbUhWbUtyVUg2bWZsNFF4NDdyRWVydjNwX3ctQldEbkN2cW4ybi1hU0tMdy1jMzlVSGtXQVVHcnEwVmxVVFh2bHk0UXBtMWxqRE5Bc3doanJTcWY5TDBCTXdpZzFJWDk1QUc0UjBTTU5ZWDU2c3NFZW9BRkZhWEd1U0pFUDNFRXpOT3ZOazQ4dHNZcTY5UTVOM3ByWUtHUnBBdg?oc=5\" target=\"_blank\">Oxford lets OpenAI train its AI models on Bodleian library</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">The Guardian</font>","full_text":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data.\nThe Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.\nOxford announced a partnership with the company in March 2025, using OpenAI software to digitise texts from the university’s world-famous library, which it said would make the content more widely available for students and researchers.\nHowever, the announcement did not state the material would be used for training OpenAI’s models, which are trained to recognise patterns in words – and thus “learn” to write complete sentences and perform other cognitive tasks – by being fed vast amounts of data.\nAn OpenAI spokesperson said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”.\n“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” they added.\nMeeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI and the effect on the university’s environmental commitments of striking a deal involving an energy-intensive technology.\nBooksellers have also reported a spate of orders for obscure titles such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners have speculated that because the titles are unlikely to exist online in digitised form, they represent fresh data that can be consumed by the next generation of AI models.\nScraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections.\nOpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project.\nBy June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other texts scanned include a rare collection of 10,000 16th-century “broadside ballads” containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitising of 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.\nThe OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.\nA spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.\nThey rejected the suggestion that the machine-learning element had been hidden from the public and students. Digitisation was the university’s primary interest, said the spokesperson, but staff had been open that the project would also contribute training data.\nThe Bodleian’s collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning. Anthropic, OpenAI’s close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books.\nA tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned.\nThe Oxford spokesperson said: “The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.\n“The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with outputs from other digitisation partnerships.\n“The project will in fact allow the Bodleian to make the material more accessible to a wider number for people, who might otherwise have found it difficult to access.”","excerpt":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.","extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 4764 characters.","diagnostics_url":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4764 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4764,"summary_length":323,"usable_text_length":4764,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4764,"summary_length":323}}},"display_formats":["compact","card","full","digest_section","json"]},"daily_stack_record":{"title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","summary":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.","source":"The Guardian","date":"2026-09-26T11:12:00+00:00","content":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data.\nThe Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.\nOxford announced a partnership with the company in March 2025, using OpenAI software to digitise texts from the university’s world-famous library, which it said would make the content more widely available for students and researchers.\nHowever, the announcement did not state the material would be used for training OpenAI’s models, which are trained to recognise patterns in words – and thus “learn” to write complete sentences and perform other cognitive tasks – by being fed vast amounts of data.\nAn OpenAI spokesperson said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”.\n“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” they added.\nMeeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI and the effect on the university’s environmental commitments of striking a deal involving an energy-intensive technology.\nBooksellers have also reported a spate of orders for obscure titles such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners have speculated that because the titles are unlikely to exist online in digitised form, they represent fresh data that can be consumed by the next generation of AI models.\nScraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections.\nOpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project.\nBy June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other texts scanned include a rare collection of 10,000 16th-century “broadside ballads” containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitising of 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.\nThe OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.\nA spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.\nThey rejected the suggestion that the machine-learning element had been hidden from the public and students. Digitisation was the university’s primary interest, said the spokesperson, but staff had been open that the project would also contribute training data.\nThe Bodleian’s collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning. Anthropic, OpenAI’s close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books.\nA tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned.\nThe Oxford spokesperson said: “The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.\n“The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with outputs from other digitisation partnerships.\n“The project will in fact allow the Bodleian to make the material more accessible to a wider number for people, who might otherwise have found it difficult to access.”","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 4764 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4764 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4764,"summary_length":323,"usable_text_length":4764,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4764,"summary_length":323}},"tags":[]},"fallback_formats":["markdown","json","html"],"actions":{"read":"/item/90757","export_markdown":"/api/items/90757/export?format=markdown","export_json":"/api/items/90757/export?format=json","diagnose":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt"},"formats":{"full":{"id":90757,"title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","source":"The Guardian","author":null,"published_at":"2026-09-26T11:12:00+00:00","locale":"en","topic":"ai","tags":[],"excerpt":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.","full_text":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data.\nThe Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.\nOxford announced a partnership with the company in March 2025, using OpenAI software to digitise texts from the university’s world-famous library, which it said would make the content more widely available for students and researchers.\nHowever, the announcement did not state the material would be used for training OpenAI’s models, which are trained to recognise patterns in words – and thus “learn” to write complete sentences and perform other cognitive tasks – by being fed vast amounts of data.\nAn OpenAI spokesperson said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”.\n“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” they added.\nMeeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI and the effect on the university’s environmental commitments of striking a deal involving an energy-intensive technology.\nBooksellers have also reported a spate of orders for obscure titles such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners have speculated that because the titles are unlikely to exist online in digitised form, they represent fresh data that can be consumed by the next generation of AI models.\nScraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections.\nOpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project.\nBy June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other texts scanned include a rare collection of 10,000 16th-century “broadside ballads” containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitising of 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.\nThe OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.\nA spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.\nThey rejected the suggestion that the machine-learning element had been hidden from the public and students. Digitisation was the university’s primary interest, said the spokesperson, but staff had been open that the project would also contribute training data.\nThe Bodleian’s collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning. Anthropic, OpenAI’s close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books.\nA tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned.\nThe Oxford spokesperson said: “The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.\n“The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with outputs from other digitisation partnerships.\n“The project will in fact allow the Bodleian to make the material more accessible to a wider number for people, who might otherwise have found it difficult to access.”","reading_time_min":4,"extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 4764 characters.","diagnostics_url":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4764 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4764,"summary_length":323,"usable_text_length":4764,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4764,"summary_length":323}}},"quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4764 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4764,"summary_length":323,"usable_text_length":4764,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4764,"summary_length":323}},"actions":{"read":"/item/90757","export_markdown":"/api/items/90757/export?format=markdown","export_json":"/api/items/90757/export?format=json","diagnose":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt"}},"digest":{"id":90757,"title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","source":"The Guardian","topic":"ai","published_at":"2026-09-26T11:12:00+00:00","excerpt":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI…","quality_bucket":"high","quality_reason":"High confidence: full text extraction produced 4764 characters.","reading_time_min":4,"cluster_id":null},"card":{"display_title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","subtitle":"The Guardian · 2026-09-26","summary":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material…","badges":["quality:high"],"links":{"read":"/item/90757","original":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","diagnose":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt"},"quality_warning":null},"export":{"title":"Oxford lets OpenAI train its AI models on Bodleian library - The Guardian","url":"https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","summary":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.","source":"The Guardian","date":"2026-09-26T11:12:00+00:00","content":"The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data.\nThe Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.\nOxford announced a partnership with the company in March 2025, using OpenAI software to digitise texts from the university’s world-famous library, which it said would make the content more widely available for students and researchers.\nHowever, the announcement did not state the material would be used for training OpenAI’s models, which are trained to recognise patterns in words – and thus “learn” to write complete sentences and perform other cognitive tasks – by being fed vast amounts of data.\nAn OpenAI spokesperson said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”.\n“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” they added.\nMeeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI and the effect on the university’s environmental commitments of striking a deal involving an energy-intensive technology.\nBooksellers have also reported a spate of orders for obscure titles such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners have speculated that because the titles are unlikely to exist online in digitised form, they represent fresh data that can be consumed by the next generation of AI models.\nScraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections.\nOpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project.\nBy June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other texts scanned include a rare collection of 10,000 16th-century “broadside ballads” containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitising of 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.\nThe OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.\nA spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.\nThey rejected the suggestion that the machine-learning element had been hidden from the public and students. Digitisation was the university’s primary interest, said the spokesperson, but staff had been open that the project would also contribute training data.\nThe Bodleian’s collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning. Anthropic, OpenAI’s close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books.\nA tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned.\nThe Oxford spokesperson said: “The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.\n“The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with outputs from other digitisation partnerships.\n“The project will in fact allow the Bodleian to make the material more accessible to a wider number for people, who might otherwise have found it difficult to access.”","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 4764 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4764 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4764,"summary_length":323,"usable_text_length":4764,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4764,"summary_length":323}},"tags":[],"format_contract_version":"news_item_formats.v1"}}}