Tuesday, March 24, 2009

Vivarium

Vivarium

Vivarium is a digital library that is a project created by the Hill Museum and Manuscript Library that has a partnership with St. John's University and the College of Saint Benedict. Vivarium specializes in Christian literature from all over the world, from medieval Europe to ancient Ethiopian manuscripts. There are several collections contained within this digital library.

The Hill Museum and Manuscript Library has a strong tradition of collecting images of monastic manuscripts and books. They originally focused on western monastic manuscripts and then gradually expanded to Eastern European and African and Indian Christian documents. Their collection guidelines specifically state that they find it a real benefit when there are photographic "back ups" of original manuscripts which has caused them to be a "library of libraries," or in other words, letting museums around the world retain guardianship of the manuscripts while Hill collects images to store in their repository to allow greater access to monastic manuscripts. With this in mind, Vivarium is pretty large and does not really have a particular selection criteria to determine eligibility of digitization.

As for metadata, each item is described through the 15 Dublin Core elements; however the metadata is not very consistent from collection to collection. For instance, the Illumination Collection's metadata is very detailed with lots of attributes, whereas the Syrian Collection just has the image and a title associated with each piece. I couldn't tell if this was because this particular collection is still a work in progress and later on, there will be more descriptive factors to aid search. When searching amongst the collections, the user has an option to search by date, author (if applicable), type, collection, letter. The list goes on; the search features are pretty detailed.


The collections within Vivarium are pretty amazing. There are a lot of different types of monastic literature featured and the digitization is done in such a way that the entirety of a piece is represented. For instance, I browsed through an Ethiopian codex. The pictures were vibrant and pretty and you could "leaf" through the codex as if you were going through the book in the physical state. This is just one of the many many many objects featured in Vivarium. Another one of my favorites was the Illumination Collection, as you could search all the manuscripts by a particular letter in the alphabet.

As for the audience, I think that Vivarium is aimed at scholars. The Ethiopian collection was started, as Vivarium states, "to stimulate scholarship on Ethiopian literature." I also think that Vivarium was created as a way for Hill Musuem and Manuscript Library to continue their objective to collect as many monastic manuscript images as possible and have a public space to showcase their collection.

Regardless, the lack of information and metadata within some of these collections irks me and I wish they would be more consistent with the metadata.

Samuel J. May Anti-Slavery Collection

Samuel J. May Anti-Slavery Collection

The Samuel J. May Anti-Slavery Collection is a Cornell digitization initiative which Cornell undertook in 1999 upon receiving a grant from NEH's Save America's Treasures program. Cornell utilized this grant to "catalog, conserve and digitize the published pamphlets [in] its Samuel J. May Anti-Slavery Collection, one of the nation's founding collections on the abolitionist movement in America" (Grant Project Description). The collection itself comprises some 10,000 pamphlets and leaflets collected by the Revered Samuel J. May during his lifetime. This author was unable to locate a precise statistic relating to how many such pamphlets and leaflets have been digitized as of today. However, all signs point towards this initiative having been completed by Cornell.

The lack of metadata with this collection came as a surprise to this author. One could perhaps argue that this is a byproduct of the collection focusing on pamphlets and leaflets. Yet, even the most basic administrative and structural metadata remains absent. The only fields one regularly finds are seemingly only the most basic of fields - title, author, collection to which the item belongs, and a date field that is viewable when browsing record lists yet nowhere to be seen when looking at an individual record but on the digitized title page of the pamphlet/leaflet itself. Thus, this author discovered no information specified by Cornell with respect to a metadata schema and its implementation. That said, Cornell does state that it employed the University of Michigan's Digital Library eXtenstion Service (DLXS) in digitizing and making available this collection (Grant Project Description). DLXS does contain a piece of middleware termed 'Collection Manager' which "maintains all collection and group information in tables" (DLXS Metadata Databases). Perhaps Cornell used this piece of middleware when employing DLXS. Then again, perhaps they did not. They do not state such information anywhere that this author accessed.

With regards to the characteristics of the digital objects, this author again had difficulty discovering a wide range of specifics. Cornell does make it known that they used a Xerox DocuImage 620S flatbed scanner for scanning and that, after scanning, they then entered a post-scanning process which included quality control and tagging amongst, presumably, other actions (Grant Project Workflow). The digital objects themselves are viewable either as digitized images of the original object or as OCR-produced text. This is a nice option to have, yet the OCR-produced text is somewhat frustrating and inaccurate. One such example is the following comparison between a random sentence fragment read from a digitized image of the original object and the same sentence fragment read from the OCR-produced text.

Digitized image of original: "But at present I shall, so far as I can, ascertain from pamphlet the specific complaints you make as to the 'emancipation proclamation' . . ."
OCR-produced text: "But at present I shall, so fhr as I can, ascertain from your pain- phiet the specific complaints you make as to the 'emancipa- tion proclamation,'. . . "

The images themselves allow for users to zoom in once while the OCR-produced text allows a user to search a specific pamphlet or the collection at large. Beyond this, this author can only say that images are downloadable as low-quality .tifs images but also printable at what appears to be a medium level of quality.

One of the pieces of this collection that this author was pleased to uncover is the following quote that Cornell used in describing the history of the Samuel J. May collection and of Cornell's Civil War collection:

In 1874 the abolitionists William Lloyd Garrison, Wendell Phillips, and Gerrit Smith, wrote, signed, and circulated an appeal to their friends and supporters in America and Great Britain, urging that it was of "great importance that the literature of the Anti-Slavery movement...be preserved and handed down, that the purposes and the spirit, the methods and the aims of the Abolitionists should be clearly known and understood by future generations. (Collection Description)

Not only does this author find this quote enjoyable from an historical perspective, this author also believes the underlying principle of Cornell's digitization attempts with respect to this collection is highlighted in this quote. That principle being to ensure that the lightnesses and the darknesses of the past never fade or disappear.

Big Orange: California Citrus Label Art


The California Historical Society created an exhibition of images of orange crate lithographs. In the 1880s, orange growers began collaborating with lithographers to create bright, eye-catching labels for their wooden crates of oranges. These lithographers began mostly making labels for wine but as the citrus industry took off in California, more lithographers began to work on designs for orange crates.

This exhibition was created in 2001 and, as a result, lacks some of the features we've seen in more recent digitization initiatives. The images are cool but there is absolutely no metadata to speak of. Thirty-five lithography companies worked on labels for orange crates but none of the images indicate which company created which image or what year they were created. As a historical society, it is not surprising perhaps that the site does a fine job of giving the historical context but does not provide data for specific images. Additionally, the images themselves are only available as thumbnails and one larger size. There is no way to zoom in further after selecting the larger sized image.

Finally, there is little information on how these particular images were chosen. It's clear that the exhibition is far from being a comprehensive sample of the lithographs made during this 70 year period but there is not explanation of why these images are digitized and not others. The historical explanation does thank three institutions for contributing so perhaps they simply digitized the materials available from these three places.

Monday, March 23, 2009

The Collier Classification System for Very Small Objects

http://www.verysmallobjects.com/html/index.html

This is a neat find, if only for the inventive and sort of whimsical classification/naming system they've created. The goal of the Collier Classification project is to make people more aware of the insignificant things around them that they might ignore as they go about their day to day lives: basically, it's a sort of art project about all things small. On to the naming!

Items in the Collier collection are given three names, each having it's own set of descriptors. An items first name is made up of two elements, one describing it's status, and the other describing it's elements. For example, a nelifrag is a small object that was never alive (neli) and a fragment of a larger whole (frag), like a pebble or a bit of chipped glass. An onliwhol, on the other hand, would be something that is complete unto itself and once living, like a dried seed. An object's second name is also made up of two descriptors: point of origin, and function. So a shosolstabscrach would be something that was found in a shoe that's apparent function was to stab, scratch, or poke things (not the best combination).

Finally, the third name of a Collier object is made up of four (!!!) descriptors: general color, general shape, consistency or surface texture, and visual comparison. So to put it all together, a Nelifrag Petfurnouse Grenirresquisunlik would be a never living fragment, found in pet fur with no apparent purpose, that was also green, irregular, squishy, and comparable to nothing but itself. Obviously!

The collection database is organized alphabetically by name, and includes a picture of the small object in question, its name, and its size. No other metadata is given, so I suppose you're supposed to pick up the naming conventions to really pick up additional information about the object. Altogether, the site seems to be an interesting artistic dissection of how and why we name and classify things, which is why I thought it would be fun to share.

Sunday, March 22, 2009

The Illuminated Books Project


The Illuminated Books Project is an independent digital collection that primarily features illustrations from books published between the 1880s and the 1920s. Even though The three individuals who are responsible for the site's content wanted to provide free access to materials that are not generally accessible to an average user due to the image's condition, rarity or location. Frustrated with their own searches for the illustrations by such artists as Walter Crane, William Morris and Kate Greenaway, they began digitizing books (their own and presumably those loaned to them) "in their own integrity and in reasonably high resolution." The images are hosted by Weatimages.

Metadata is primarily given in a book's introduction - known information about original publications, exhibitions, and relative biographical information on the author/illustrator. There is no data on current ownership of the book, loan information, date of digitization, scanning details, or cross-references. Users are given the options to view the JPEG images in different sizes and a one-click enhancement can give a close-up. According to the "Using the Website" page, the "intention is to provide high quality images with no heavy compression artifacts. Most images are available at the original size of the page (or spread) and at 150dpi for each full-sized JPEG/JPG image. File sizes with resolutions of 150dpi can be quite large and time-consuming to download for those visiting with low speed modem connections. In the event that a book is presented at a resolution higher than 150 dpi, the size will be posted in the page of the specific book."
The project includes both adult and children's books. Books are organized by author/illustrator without any keyword or Boolean search options. Each book is digitized cover to cover with double or single-page spread thumbnail images to click on. There are no user functions for developing individual metadata or personal collections. So even though the site's intent is to feature these decorative images, a user will have to contend with pages of text while searching for an image. It is hard to tell when the site was last updated or if it is still being maintained. The most current date I could locate was 2006. The Spanish language option is still under construction.

Although the collection is small compared to those hosted by universities and there are no search options to locate specific images, I think the images are digitized well. The site designers want people to be able to look at the images closely, and they made that possible. It is obviously geared toward an audience who is interested in highly decorative illustrations and willing to browse through entire texts to find them. A researcher may find the illustrations useful but probably not for academic research.

Thursday, March 19, 2009

Women Working, 1800-1930

http://ocp.hul.harvard.edu/ww/

Women Working is the first collection in Harvard University Library’s Open Collections Program (OCP). The OCP is a seemingly well-planned initiative, which means there’s interesting documentation about the collection decisions. The OCP explains briefly that its selection standards are “to create comprehensive, subject-based digital collections through the careful selection of topics and materials,” and the Women Working collection narrative expands on that overall standard to offer specific criteria and guidelines for selection. The criteria for subject selection are focused on utilizing a range of Harvard Libraries, usefulness in teaching, range of kinds of objects, specificity, not duplicating the work of other projects, and willingness of Harvard faculty to engage with the topic development. Selection guidelines assert that materials for Working Women should be “academically significant,” should “reflect women’s experience working,” should cover controversies over their full range of discussion, should “appeal to younger students,” and should draw on a range of types of material. Further, the collection narrative explains the steps researchers took to identify materials, including online catalog searches, shelf browsing, and review of bibliographies and studies related to the topic. Finally, the collection indicates that information architects examined the physical condition of the materials, the usefulness of the visual object, and the translatability from the page to the screen (according to cost and navigability). Items are duplicated in certain situations, including when an item is only available commercially; this is an interesting feature considering that it is nicely in line with the topic of Women Working and attention to access.

The collection went live in 2004 and digitization continued through 2005; the collection includes books, pamphlets, serials, consumer and trade catalogs, magazines, photographs, and manuscripts including diaries.

All objects include metadata for title, creator, and date on the browsing list. The titles are descriptive. There is a link for full display which includes extensive metadata for object location, place of origin, publisher, language, description of size of the original object, form/genre, subject, categories, notes, and other titles. It is not clear whether these metadata were automatically generated (or partially automatically generated) from existing materials, since some of the objects are linked to existing collections. It appears from the uniformity of available information that at least some of the metadata is original or draws on existing Harvard Library standards. I am particularly pleased to see the metadata category for publisher, since this has come up with our own digitization project.

The digital objects are JPEGs with RGB color. The project outlines directions for downloading and printing images, so I assume that the resolution is higher than web quality. I didn’t find a description of the digitization process, which was disappointing.

The audience is primarily teachers; the fourth link on the menu bar is Teacher Resources, which include five lesson themes with extensive materials including images, texts, and more. The themes challenge viewers by demonstrating how the materials allow us to understand the construction of our understandings of race, national identity, childhood, capitalism, and gendered work with lessons like “‘The Materials of the New Race’: Immigration and Whiteness” and “Soap and Settlements: ‘Making a Cleaner Society.’”

Sunday, March 15, 2009


The Szathmary Recipe Pamphlet Digital Collection is, compared to some of the collections I've waded through, superb. The collection itself was gathered by one Louis Szathmary, a Chicago restauranteer, and a portion of it is now held at the University of Iowa Libraries. We've looked at collections from this library system before, as I recognized parts of the interface.

The objects themselves are "representative samples" from 1880 through 1930. Though it is not explicitly stated as such, the intro blurb on the main page mentions how this era showed a particular shift in eating habits among americans, and that this change is reflected in the ephemera gathered here, so I'd take that as a collection policy. Many characteristics of the objects and collection are covered in the metadata, and with an included "reference URL" link on each object page it appears to be fairly interoperable. One complaint is the zoom function: it is integrated into the object's page, which is nice, but must reload every time you want to move to the left or right, up or down, or zoom in, a tedious operation for someone visually scanning the object up close.

The metadata provided is ample, and easily outshines any collection I've covered thus far. Fields are included not just for the standard Title, Publisher, and Date, but also for all sorts of subject headings (conforming to LCSH, DCMITV and a couple of other acronyms I didn't recognize) and fields including digital collection, contributing collection, and archival collection, as well as Rights Management, Contact, Digitization Specifications and Date Digital. This is a wealth of information, and most terms used in these fields are searchable, i.e. you can click on them and it will bring you all the objects which match that particular criteria. I would say the metadata conforms fully to the guidelines laid out by the NISO framework.

As for the objects themselves, each has an archival quality .TIF image available, and the authenticity can be fairly easily deduced by the substantial metadata. The images are broadly accessible (though they failed the google image search) but are of a very, very large size, as you can plainly see, good for research but not so good for interoperability or use.

I think the target audience of this would be scholars. the stated collection policy mentions what this ephemera reflects, and it's not hard to see a scholar using the robust search capability and zooming ability to easily compare and contrast objects.