Showing posts with label digital library. Show all posts
Showing posts with label digital library. Show all posts

Friday, January 1, 2010

SALT Project lessons learned

 I recently answered a kind inquiry about the lessons learned during the 3-yr fixed term SALT Project at Stanford.  I thought my informal answers were worth posting here.

1) The foundation of SALT, the database, was very similar to DSpace.  When I arrived on the project, SALT was using an in-house developed database that the Stanford Libraries had created before DSpace became popular.  After the completion of our SALT Prototype, the SALT database is being ported to Fedora.  So open-source tools are viewed as very desirable.  What made SALT unique and different was the user "access" side: the ability to enable and deliver visually-oriented views into the material.  We showed information summaries in "graph network" maps and timelines, and then the user could drill down into a faceted browser to search for individual objects (papers) of interest.  We used Berkeley's Flamenco for the faceted browser, later replaced by UVAs Blacklight.  But the visual representations of the faculty legacy was the thing that was the most unique and drew the faculty's interest.

2) The SALT Project was a collaboration between the Stanford Libraries and a professor emeritus of Computer Science.  So the project had a flavor of research, even though the intent was not to develop a lot of software in-house but to package open source solutions to achieve the documented requirements.  The team was asked to apply for an NSF grant, which we did, so the objectives had a strong flavor of research.  As you can imagine, the requirements were fairly visionary and it was a challenge to deliver working examples of the elements.  However, we mostly tried to find existing solutions as I mentioned above ... and for each solution chosen we evaluated several.  We implemented the SALT Prototype in 3 iterations.  With each prototype release, we got a better understanding of the goals and requirements of the various stakeholders (faculty, archivists, engineers, curators, researchers, management).  In our final version, all the major software functions were off-the-shelf and the software we developed just connected the elements of the tool chain together.

3) The faculty reaction has been very positive.  I interviewed a handful before and after delivering our prototype, and they were very interested in documenting their legacy and especially tracking the flow of their teaching through their students.  The "graph network" map was the key to drawing them in to offer "the backstory" and new knowledge.   They were most willing to offer information verbally, rather than typing (the older faculty were challenged by eyesight and typing).  They particularly valued that the system was largely automatic, because they are very limited on time to enter data or insights. 

4) SALT is moving forward as a "production" project within the Stanford Digital Library group, mostly taking steps to integrate the object-level documents into the library's standard catalog.  I do indeed anticipate some problems still to be solved, for example with the dis-ambiguation of names and tag metadata.  There is not a good point-solution for this problem, but Stanford's archival staff were looking to the Library of Congress classifications.  There needs to be more thought and integration with traditional archival data-views, Stanford's archival staff and the researchers using the collections were more comfortable with finding aids in EAD format; I think these could be easily generated as an alternate view.  Ultimately the biggest problem is the volume of faculty archival information far exceeding the archival staff's ability to process it, so all stakeholders will need become more comfortable with automated solutions.  Lastly, there is an unsolved problem with copyright and privacy liability vs offering access, especially with regard to email.  Email is becoming more and more a way for archiving and collaboration.  See also: papers by Cathy Marshall of Microsoft. 

I continue to think that archiving of faculty papers is a very important area, as faculty papers are more and more born-digital an faculty are finding solutions for self-archiving.   I found that the cutting edge profs were already getting their staffs to digitize important papers and lab notebooks, and create online archives with self-description.  I think it will serve the University and patrons to help the faculty be successful with these efforts, and in the process to organize their material in a way that is more easily ingested into the Library's preservation system.

Saturday, November 1, 2008

Google strikes a deal with publishers to make more books available online

from The Economist:

Google has reached a settlement with groups representing authors and publishers, who had sued the world’s biggest internet company in 2005 for copyright infringement. Google has been scanning millions of books into digital form and making the text available online in small “snippets”, to answer users’ search queries. The company said this constituted “fair use” of the copyrighted material, since only a small part of the work was accessible. The Association of American Publishers and the Authors Guild disagreed. The resulting legal fight was widely expected to end in a settlement, but it took a while to agree on the terms.

Under the deal, Google will pay $125m to compensate copyright holders and will set up a “Book Rights Registry” to pay authors and publishers for the digital use and sale of works—akin to the collection societies that send royalty cheques to musicians and songwriters.

http://www.economist.com/business/displaystory.cfm?story_id=12523914

Sunday, May 25, 2008

Stephen Jay Gould Collection arrives at Stanford


During his career, Gould wrote 300 consecutive essays for Natural History, the monthly magazine of the American Museum of Natural History, and more than 20 books, many of them bestsellers. He also assembled what he believed was a definitive library of the history of early paleontology, said Rhonda Shearer, Gould's widow.

Now, the collection of books, papers and artifacts that helped inform his writing and teaching is, for the most part, in the Stanford University Libraries, with the balance expected to arrive soon. It is an immense amount of material.

...

"Stanford was the only institution really prepared to make a commitment to digitize and cross-link all of Steve's work, and this is something that Steve wanted," said Shearer. "Even though he called himself a Luddite and really had anxiety about technology, he saw that for ideas to compete, they really had to be out on the Internet."

[Michael] Keller said the plan is to digitize Gould's articles, as well as the sources from which he drew both inspiration and information, and cross-link the source materials to the endnotes and citations in his writing. The goal will be to make all of Gould's papers freely available over the Internet to anyone who wants to see them, whether schoolchildren or scholars.

"From very detailed explorations with other institutions, Stanford was head and shoulders above all the rest in its ability to fulfill such a promise, as well as be in a technical position of expertise with the ability to execute," Shearer said.

...

"We see [this project] as a kind of model of what could be done with the emanations of really brilliant scientists, and we certainly have that same ambition to work on many of Stanford's own leading thinkers in the sciences and across all the disciplines," Keller said.

http://news-service.stanford.edu/news/2008/may14/gould-051408.html

Tuesday, April 29, 2008

SALT at DLF

I presented a "paper" on the Self Archiving Legacy Toolkit (SALT) Project at the Digital Library Federation Meeting in Minneapolis, April 29, 2008.

Abstract

This paper describes Stanford's work on the Self-Archiving Legacy Toolkit (SALT), which is exploring the capability of a human-computer system to transform unstructured and heterogeneous data into navigable collections of information presented in context. Stanford's University Archives are beginning to digitize and present the collected "papers" of luminary faculty. The full potential of this resource of scientific legacy is not realized by simple digitization. Through the application of semantic processing technologies (coupled with rich visualization tools), a luminary's lifetime collection of research, publications, correspondence and presentations can be accessed not only by keyword, but also by concept, collaborators, time, place, organization and even project. Our hypothesis is that these facets will transform the processing and delivery of personal archival collections, and expose the historical context and intellectual concepts threading through the careers of some of the greatest scientists and thinkers of the 20th century.

Central to the vision of SALT is the notion of self-description of a luminary's own corpus. In addition to providing oral history, video commentary and textual annotation to their collected works, the toolkit gives eminent researchers the tools to create, apply and edit their own taxonomies, ontologies and controlled vocabularies to their works. These knowledge editing tools will enable the archival subject to efficiently identify people, places, times, and concepts mentioned in their collected works and materials, and to draw relationships among these items to reveal paths of influence and the historic progression of ideas. In this way, they can interpret and extend their collection with their own personal viewpoint, creating a uniquely personal presentation of their own life story, which augments rather than replaces the traditional privileged view of archival provenance.

link to the presentation in pdf

Tuesday, May 29, 2007

Jeff Rothenberg of RAND

Jeff Rothenberg of RAND visited Friday to discuss preservation and emulation of digital information.

A first interesting thought: what is a "digital original?" There will be the original and then the surrogate, which may receive transliteration to be current to the context of the reader ... a surrogate copy. New surrogates would be created from the original, and the original preserves trust in the surrogates.

Any digital object is an executable file, designed to be interpreted. An important point was that "born digital" documents are those which are equivalent to text typed on a page, where "inherently digital" objects need to be rendered by a computer. The vast bulk of digital information is not inherently digital, but we are generating more inherently digital information than we expected. However, a born-digital document may receive value-added in the form of comments, links, labels, annotations; and these are "inherently digital" and are value that needs preservation.

Inherently digital information will need emulation to interpret it over time. Emulation, migration, formalization steps.

Key take-away: we know very well how to preserve paper documents; digitization priority is therefore not driven by preservation. Digitization is driven by access, re-use, the ability for multiple users to create inherently digital works of value.


Jeff Rothnberg "Avoiding Technological Quicksand" ...
http://www.clir.org/PUBS/reports/rothenberg/contents.html