Thursday, September 4, 2008

Navigating the network of knowledge: Mining quotations from massive-scale digital libraries of books

a lecture at PARC

Bill Schilit, Google Research

Scanning books, magazines, and newspapers is widespread because people believe a great deal of the world's information still resides off-line. In general, after works are scanned they are indexed for search and processed to add links. In this talk I will describe a new approach to automatically add links by mining repeated passages. This technique connects elements that are semantically rich, so strong relations are made. Moreover, link targets point within rather than to the entire work, facilitating navigation. Our system has been run on a digital library of over 1 million books (Google Book Search), has been used by thousands of people, and has generated the world's largest collection of quotations. I will also present a follow-on project based on the theory that authors copy passages from book to book because these quotations capture an idea particularly well: Jefferson on liberty; Stanton on women’s rights; and Gibson on cyberpunk. These projects suggest that mining quotations for links and ideas are an important mechanism for understanding the knowledge contained in books.

video archived at: http://www.parc.com/cms/get_article.php?id=772

------------

The automatic creation of the idea map is very interesting. These algorithms may be used effectively on the collected papers of luminaries as well as books. Papers leading up to ideas expressed in books may be linked into the idea map and easily found. The luminary themselves could add additional links and meaning.

Friday, August 29, 2008

SALT at SAA

I was invited to give my talk on the SALT Project to the Society of American Archivists (Science, Technology, and HealthCare Roundtable) by Paul Theerman of the National Library of Medicine, chair. This was at their annual meeting, San Francisco, August 27, 2008. My thanks to the Roundtable for the warm reception, good questions, and great ideas for collaborations.


Sunday, May 25, 2008

Stephen Jay Gould Collection arrives at Stanford


During his career, Gould wrote 300 consecutive essays for Natural History, the monthly magazine of the American Museum of Natural History, and more than 20 books, many of them bestsellers. He also assembled what he believed was a definitive library of the history of early paleontology, said Rhonda Shearer, Gould's widow.

Now, the collection of books, papers and artifacts that helped inform his writing and teaching is, for the most part, in the Stanford University Libraries, with the balance expected to arrive soon. It is an immense amount of material.

...

"Stanford was the only institution really prepared to make a commitment to digitize and cross-link all of Steve's work, and this is something that Steve wanted," said Shearer. "Even though he called himself a Luddite and really had anxiety about technology, he saw that for ideas to compete, they really had to be out on the Internet."

[Michael] Keller said the plan is to digitize Gould's articles, as well as the sources from which he drew both inspiration and information, and cross-link the source materials to the endnotes and citations in his writing. The goal will be to make all of Gould's papers freely available over the Internet to anyone who wants to see them, whether schoolchildren or scholars.

"From very detailed explorations with other institutions, Stanford was head and shoulders above all the rest in its ability to fulfill such a promise, as well as be in a technical position of expertise with the ability to execute," Shearer said.

...

"We see [this project] as a kind of model of what could be done with the emanations of really brilliant scientists, and we certainly have that same ambition to work on many of Stanford's own leading thinkers in the sciences and across all the disciplines," Keller said.

http://news-service.stanford.edu/news/2008/may14/gould-051408.html

Tuesday, April 29, 2008

SALT at DLF

I presented a "paper" on the Self Archiving Legacy Toolkit (SALT) Project at the Digital Library Federation Meeting in Minneapolis, April 29, 2008.

Abstract

This paper describes Stanford's work on the Self-Archiving Legacy Toolkit (SALT), which is exploring the capability of a human-computer system to transform unstructured and heterogeneous data into navigable collections of information presented in context. Stanford's University Archives are beginning to digitize and present the collected "papers" of luminary faculty. The full potential of this resource of scientific legacy is not realized by simple digitization. Through the application of semantic processing technologies (coupled with rich visualization tools), a luminary's lifetime collection of research, publications, correspondence and presentations can be accessed not only by keyword, but also by concept, collaborators, time, place, organization and even project. Our hypothesis is that these facets will transform the processing and delivery of personal archival collections, and expose the historical context and intellectual concepts threading through the careers of some of the greatest scientists and thinkers of the 20th century.

Central to the vision of SALT is the notion of self-description of a luminary's own corpus. In addition to providing oral history, video commentary and textual annotation to their collected works, the toolkit gives eminent researchers the tools to create, apply and edit their own taxonomies, ontologies and controlled vocabularies to their works. These knowledge editing tools will enable the archival subject to efficiently identify people, places, times, and concepts mentioned in their collected works and materials, and to draw relationships among these items to reveal paths of influence and the historic progression of ideas. In this way, they can interpret and extend their collection with their own personal viewpoint, creating a uniquely personal presentation of their own life story, which augments rather than replaces the traditional privileged view of archival provenance.

link to the presentation in pdf

Thursday, April 17, 2008

Freebase: An Open Database of the World's Information

a speech at PARC:

John Giannandrea, Metaweb Technologies

Freebase -- an open database of the world's information -- is built by a global community and is free for anyone to query, contribute to, and build applications on. Drawing from large open data sets like Wikipedia, MusicBrainz, and the SEC, Freebase is curated by a passionate global community of users and contains structured information on millions of topics such as music, food ingredients, rocket engines, stock indices, historical events, and more.

Part of what makes this open database unique is that it spans domains, but requires that a particular topic exist only once in Freebase -- even if it might normally be found in multiple databases. For example, Arnold Schwarzenegger would appear in a movie database as an actor, a political database as a governor, and in a bodybuilder database as a Mr. Universe. In Freebase, however, there is only one topic for Arnold Schwarzenegger that brings all these facets together. The unified topic is a single reconciled identity, which makes it easier to find and contribute information about the linked world we live in.

archive video at:
http://www.parc.com/cms/get_article.php?id=736

------------

Freebase is designed so software may interact more easily with the data. There are may interesting and useful applications that use Freebase data, making Freebase more of "switchboard" of data than an end in itself (a "graph data store"). Freebase may be very useful as a "name authority," a unique web location to define a person, place, organization, concept. It can also store relationships of these entities as "triples." There phiosophy of "Schema Later" seems to apply well to an evolving luminary archive.

Tuesday, May 29, 2007

Semantic Web Conference 2007

http://www.semantic-conference.com/

Key take-aways
  • Semantic web technologies are maturing into businesses
  • Semantic + social solutions is a buzz
  • Companies realizing the value of internal semantic search
  • Federal grant spending is reaching it's 5-yr objectives, ready to commercialize.

Jeff Rothenberg of RAND

Jeff Rothenberg of RAND visited Friday to discuss preservation and emulation of digital information.

A first interesting thought: what is a "digital original?" There will be the original and then the surrogate, which may receive transliteration to be current to the context of the reader ... a surrogate copy. New surrogates would be created from the original, and the original preserves trust in the surrogates.

Any digital object is an executable file, designed to be interpreted. An important point was that "born digital" documents are those which are equivalent to text typed on a page, where "inherently digital" objects need to be rendered by a computer. The vast bulk of digital information is not inherently digital, but we are generating more inherently digital information than we expected. However, a born-digital document may receive value-added in the form of comments, links, labels, annotations; and these are "inherently digital" and are value that needs preservation.

Inherently digital information will need emulation to interpret it over time. Emulation, migration, formalization steps.

Key take-away: we know very well how to preserve paper documents; digitization priority is therefore not driven by preservation. Digitization is driven by access, re-use, the ability for multiple users to create inherently digital works of value.


Jeff Rothnberg "Avoiding Technological Quicksand" ...
http://www.clir.org/PUBS/reports/rothenberg/contents.html