I recently answered a kind inquiry about the lessons learned during the 3-yr fixed term SALT Project at Stanford. I thought my informal answers were worth posting here.
1) The foundation of SALT, the database, was very similar to DSpace. When I arrived on the project, SALT was using an in-house developed database that the Stanford Libraries had created before DSpace became popular. After the completion of our SALT Prototype, the SALT database is being ported to Fedora. So open-source tools are viewed as very desirable. What made SALT unique and different was the user "access" side: the ability to enable and deliver visually-oriented views into the material. We showed information summaries in "graph network" maps and timelines, and then the user could drill down into a faceted browser to search for individual objects (papers) of interest. We used Berkeley's Flamenco for the faceted browser, later replaced by UVAs Blacklight. But the visual representations of the faculty legacy was the thing that was the most unique and drew the faculty's interest.
2) The SALT Project was a collaboration between the Stanford Libraries and a professor emeritus of Computer Science. So the project had a flavor of research, even though the intent was not to develop a lot of software in-house but to package open source solutions to achieve the documented requirements. The team was asked to apply for an NSF grant, which we did, so the objectives had a strong flavor of research. As you can imagine, the requirements were fairly visionary and it was a challenge to deliver working examples of the elements. However, we mostly tried to find existing solutions as I mentioned above ... and for each solution chosen we evaluated several. We implemented the SALT Prototype in 3 iterations. With each prototype release, we got a better understanding of the goals and requirements of the various stakeholders (faculty, archivists, engineers, curators, researchers, management). In our final version, all the major software functions were off-the-shelf and the software we developed just connected the elements of the tool chain together.
3) The faculty reaction has been very positive. I interviewed a handful before and after delivering our prototype, and they were very interested in documenting their legacy and especially tracking the flow of their teaching through their students. The "graph network" map was the key to drawing them in to offer "the backstory" and new knowledge. They were most willing to offer information verbally, rather than typing (the older faculty were challenged by eyesight and typing). They particularly valued that the system was largely automatic, because they are very limited on time to enter data or insights.
4) SALT is moving forward as a "production" project within the Stanford Digital Library group, mostly taking steps to integrate the object-level documents into the library's standard catalog. I do indeed anticipate some problems still to be solved, for example with the dis-ambiguation of names and tag metadata. There is not a good point-solution for this problem, but Stanford's archival staff were looking to the Library of Congress classifications. There needs to be more thought and integration with traditional archival data-views, Stanford's archival staff and the researchers using the collections were more comfortable with finding aids in EAD format; I think these could be easily generated as an alternate view. Ultimately the biggest problem is the volume of faculty archival information far exceeding the archival staff's ability to process it, so all stakeholders will need become more comfortable with automated solutions. Lastly, there is an unsolved problem with copyright and privacy liability vs offering access, especially with regard to email. Email is becoming more and more a way for archiving and collaboration. See also: papers by Cathy Marshall of Microsoft.
I continue to think that archiving of faculty papers is a very important area, as faculty papers are more and more born-digital an faculty are finding solutions for self-archiving. I found that the cutting edge profs were already getting their staffs to digitize important papers and lab notebooks, and create online archives with self-description. I think it will serve the University and patrons to help the faculty be successful with these efforts, and in the process to organize their material in a way that is more easily ingested into the Library's preservation system.
Friday, January 1, 2010
Friday, August 28, 2009
Moving On
The Stanford SALT Project has since April been transitioning to the DLSS production development team. The start-up phase of the SALT Project is complete and I will be leaving SULAIR at the end of the fiscal year.
SALT in production will be built on the foundation being developed for Stanford Searchworks. This is a significant positive step toward realizing the value of digital online archives of distinguished faculty for historical research.
While I am moving into a new project, I am very interested to keep continuity and momentum of SALT activities and will be glad to continue to help, or to guide you to the appropriate SULAIR resource.
SALT in production will be built on the foundation being developed for Stanford Searchworks. This is a significant positive step toward realizing the value of digital online archives of distinguished faculty for historical research.
While I am moving into a new project, I am very interested to keep continuity and momentum of SALT activities and will be glad to continue to help, or to guide you to the appropriate SULAIR resource.
Monday, August 10, 2009
SALT Demo
I just put a link to our demo on the SALT website luminaryarchives.stanford.edu. This links to our "mash up" prototype. It is live an click-able, except you will not be able to see actual documents unless you have a Stanford ID and our team gives you authorization.
Wednesday, June 10, 2009
SALT Background (video)
I've posted a short video about the SALT Project on our site luminaryarchives.stanford.edu. This was a collaborative effort of our team and especially our Operations Assistant Rachael. I think it came out very well, please see it at this direct link to the SALT Movie.
Tuesday, February 10, 2009
SALT at British Library
I presented the SALT Project at the Digital Lives conference in London. The area of archiving, hosting, and presenting personal papers in the digital age is interesting and going to be of increasing importance. I've posted the current version of the slides on slideshare and linked to the SALT site.
Saturday, November 1, 2008
Google strikes a deal with publishers to make more books available online
from The Economist:
Google has reached a settlement with groups representing authors and publishers, who had sued the world’s biggest internet company in 2005 for copyright infringement. Google has been scanning millions of books into digital form and making the text available online in small “snippets”, to answer users’ search queries. The company said this constituted “fair use” of the copyrighted material, since only a small part of the work was accessible. The Association of American Publishers and the Authors Guild disagreed. The resulting legal fight was widely expected to end in a settlement, but it took a while to agree on the terms.
Under the deal, Google will pay $125m to compensate copyright holders and will set up a “Book Rights Registry” to pay authors and publishers for the digital use and sale of works—akin to the collection societies that send royalty cheques to musicians and songwriters.
http://www.economist.com/business/displaystory.cfm?story_id=12523914
Google has reached a settlement with groups representing authors and publishers, who had sued the world’s biggest internet company in 2005 for copyright infringement. Google has been scanning millions of books into digital form and making the text available online in small “snippets”, to answer users’ search queries. The company said this constituted “fair use” of the copyrighted material, since only a small part of the work was accessible. The Association of American Publishers and the Authors Guild disagreed. The resulting legal fight was widely expected to end in a settlement, but it took a while to agree on the terms.
Under the deal, Google will pay $125m to compensate copyright holders and will set up a “Book Rights Registry” to pay authors and publishers for the digital use and sale of works—akin to the collection societies that send royalty cheques to musicians and songwriters.
http://www.economist.com/business/displaystory.cfm?story_id=12523914
Thursday, September 4, 2008
Navigating the network of knowledge: Mining quotations from massive-scale digital libraries of books
a lecture at PARC
Bill Schilit, Google Research
Scanning books, magazines, and newspapers is widespread because people believe a great deal of the world's information still resides off-line. In general, after works are scanned they are indexed for search and processed to add links. In this talk I will describe a new approach to automatically add links by mining repeated passages. This technique connects elements that are semantically rich, so strong relations are made. Moreover, link targets point within rather than to the entire work, facilitating navigation. Our system has been run on a digital library of over 1 million books (Google Book Search), has been used by thousands of people, and has generated the world's largest collection of quotations. I will also present a follow-on project based on the theory that authors copy passages from book to book because these quotations capture an idea particularly well: Jefferson on liberty; Stanton on women’s rights; and Gibson on cyberpunk. These projects suggest that mining quotations for links and ideas are an important mechanism for understanding the knowledge contained in books.
video archived at: http://www.parc.com/cms/get_article.php?id=772
The automatic creation of the idea map is very interesting. These algorithms may be used effectively on the collected papers of luminaries as well as books. Papers leading up to ideas expressed in books may be linked into the idea map and easily found. The luminary themselves could add additional links and meaning.
Bill Schilit, Google Research
Scanning books, magazines, and newspapers is widespread because people believe a great deal of the world's information still resides off-line. In general, after works are scanned they are indexed for search and processed to add links. In this talk I will describe a new approach to automatically add links by mining repeated passages. This technique connects elements that are semantically rich, so strong relations are made. Moreover, link targets point within rather than to the entire work, facilitating navigation. Our system has been run on a digital library of over 1 million books (Google Book Search), has been used by thousands of people, and has generated the world's largest collection of quotations. I will also present a follow-on project based on the theory that authors copy passages from book to book because these quotations capture an idea particularly well: Jefferson on liberty; Stanton on women’s rights; and Gibson on cyberpunk. These projects suggest that mining quotations for links and ideas are an important mechanism for understanding the knowledge contained in books.
video archived at: http://www.parc.com/cms/get_article.php?id=772
------------
The automatic creation of the idea map is very interesting. These algorithms may be used effectively on the collected papers of luminaries as well as books. Papers leading up to ideas expressed in books may be linked into the idea map and easily found. The luminary themselves could add additional links and meaning.
Subscribe to:
Posts (Atom)