All posts by admin

Crowdsourcing the Digital Humanities, Part II

Almost every definition of the Digital Humanities includes that it is a collaborative undertaking. Collaboration is not just with peers, but includes engaging the public in crowdsourcing. Wikipedia is the most visible of the crowdsourcing projects,

We have looked at AI, Wikipedia and transcription projects as examples of crowdsourcing. I prefer transcription for crowdsourcing. Transcription engages the public in a different way than a Wikipedia type project, in that transcription can connect the volunteer on many levels. First, they become part of something larger, and can contribute to spreading knowledge. Crowdsourcing the tasks of transcription brings a lot of potential manpower to a project, but also fosters connection with a collection and its goals. It can also connect a transcriber to history on a more personal level. A volunteer, with good instruction can transcribe and review and engage with artifacts in a completely different way than they are accustomed to. Volunteers can also help with determining the allocation of resources, as projects that more people are interested in working on can also help understand what the public is interested in and inform other decision.

Crowdsourced projects demand engaged participants and retaining participants to meet goals for work accomplishment. Engagement is related to value – making suer the participants feel valued and presenting the project in a way that demonstrated value. Shakespeare’s World and By the People both do this is several ways. While Shakespeare’s World his currently on hiatus for data processing, exploring the site shows some of the ways participant interest can be maintained. The Talk section shows some of the ways participants are engaged. The page includes news, FAQs, help with handwriting, areas of interest to researchers, OED (the project has contributed entries) and Recipes. Recipes invites the participants to try the ones they find and add results for the community.

#Recipes2Try, https://www.zooniverse.org/projects/zooniverse/shakespeares-world/talk/228

Efforts with the researchers and volunteers resulted in working on understanding a recipe with a contestant from the Great British Baking Show, who published on it, and an entry into the OED.1

In By the People, value of the participants in their work comes in several ways. First, the instructions are incredibly user friendly, making it a welcoming site. Using the Car barton Campaign as an example, choosing an item included a comment that the volunteer may be the first one in 100 years to be reading her letterbook — the volunteer is now part of a select group. The project goes on to encourage participants as they grow more confidant, to add doing peer reviews to what they are doing, again recognizing that they are key to not only transcription, but their increasing expertise is trusted to peer review the work.

  1. Van Hyning, Victoria Anne, and Mason A. Jones. “Data’s Destinations: Three Case Studies in Crowdsourced Transcription Data Management and Dissemination.” Startwords 2 (2021)., pp 5-7.  https://startwords.cdh.princeton.edu/issues/2/datas-destinations/
    Links to an external site.
    .

    PreviousNext







    ↩︎

Crowdsourcing

WIkipedia’s reputation has improved from its inception, a site with the radical idea that non-experts should be able to contribute to knowledge on the world wide web. Not long ago, I would not consider Wikipedia a valid resource, but while I would not use it as a main source, I have found it is often a jumping off point for research. Understanding how Wikipedia entries is important in evaluating the value of an article and its accuracy. Given that chat bots also use it as one of their large language models, the ability to review entries and how they develop is useful.

Using the Digital Humanities page as an example, the “Talk” and “View history” sections provide insight into the development of the page and its contents. “View history” allows the viewer to see how the site has changed, as well as a side by side comparison of what each change did to the site. It also identifies the editor, linking to their user page and frequently a description of their expertise. 

Digital humanities: Revision history at https://en.wikipedia.org/w/index.php?title=Digital_humanities&action=history

In the screenshot above, examples include a 27 November 2022 edit for unsourced material being added; a 3 June 2022 addition on Analysis of macroscopic trends in cultural change was added, then changed slightly.; and on 5 January 2022 entries reflecting the rendition of the unsourced work with comment, and the subsequent removal of it again. The 27 November 2022 change is made by ElKevbo, the third highest contributor to the site, and/or his biography, works in the digital humanities and higher learning.

Adding to this are comments in the “Talk” section, a discussion of what content should be added. I particularly liked the following post, discussing areas to improve their understanding and that of other users:

Technology, criticisms, & future at https://en.wikipedia.org/wiki/Talk:Digital_humanities

The section also discusses some of the editing decisions and collaboration, such as an early crowdsourcing event at THATcamp (The Humanities and Technology Camp) in Australia making several edits. The first large batch of edits were made by the in the field thinking and learning about the digital humanities.

Understanding the source of the information, and that Wikipedia editors are trying to maintain community standards, lends credibility to the entries. AI crowdsources information through the text it was trained on, including Wikipedia. While it synthesizes information rapidly, and by using large, internet based data files, it is still only as good as the crowd it is using. Its inability to show sources makes it another useful starting point, but not necessarily an ending point. I queried Chat GPT about its use of Wikipedia, with this response:

I don’t have direct access to the internet or specific databases like Wikipedia in real-time. Instead, I’ve been trained on a diverse range of internet text available up until my last update in 2022, which includes information from sources like Wikipedia. My responses are generated based on patterns and information learned from that data.

If you ask me a question about a topic, I’ll generate a response based on the knowledge I’ve acquired during my training. So, while I can’t access Wikipedia directly, I’ve learned from its content and can provide information based on what I’ve previously read and absorbed.

Wikipedia, a world of words askew,
A dance of data, a linguistic brew.
No capital chains to bind its grace,
In ee cummings’ style, it finds its place.

A symphony of facts in a tangled rhyme,
Hyperlinks weaving through the sands of time.
No rigid structure, just free-flowing streams,
In Wikipedia’s garden, knowledge dreams.

Lowercase whispers tell tales untold,
As wikis unfold, in pixels bold.
Syntax surrenders, conventions flee,
In this textual playground, wild and free.

Beneath the letters, a universe blooms,
A kaleidoscope of knowledge looms.
No need for order, no need for rule,
Wikipedia dances in the vastness, cool.

A symphony of voices, a chorus untamed,
In cummings’ spirit, where no word’s framed.
Wikipedia, a poem of pixels and links,
In the language of chaos, information syncs.

wiki leaks words, a
cacophony, a
melange of info
sprawling unchained un
bridled in its verve un
stop-pable unwavering it grows

words dance in links
a tap dance of thoughts
meandering through time space
no capitals or spaces
just the pulse of info
in a torrential race

oh, wikipedia!
encyclopedic cacophony
where knowledge meets chaos
in a glorious blend
of facts and conjecture
yet a haven for seekers
in a digital splendor un
matched and profound

The poems above illustrate both the questionable and good about AI crowdsourcing. The Chat GPT prompt was “write a poem in the style of ee cummings.”The first try has the syntax and rhythm of Cummings, but it was not until the second try that it picked up the trait Cummings is known for, the lack of capitalization. AI did not recognize initially, although the Wikipedia entry on Cummings describes it, with several examples of poems.

The challenge of crowdsourcing, through a Wikipedia or Wikipedia of the future or AI, remains the expertise and the biases of the crowd. Lockett engagement with her students at Spelman as being motivated by realizing how many gaps in information existed about their college and those like it, and Black women in particular1. Crowdsourcing brings a myriad of experiences, expertise and interests to the world in a very shareable way, however moving into the future, identifying the gaps and blind spots will be critical to their overall effectiveness as a tools.

  1. Lockett, Alexandria. “Why Do I Have Authority to Edit the Page? The Politics of User Agency and Participation on Wikipedia.” In Wikipedia @ 20: Stories of an Incomplete Revolution. Edited by Joseph Reagle and Jackie Koerner (MIT Press,
    2020), https://doi.org/10.7551/mitpress/12366.003.0019Links to an external site.., p 213. ↩︎

Digital Tools Compared

Working with three different digital tools highlighted the ways digital tools can help understand data as well as develop new lines of research.  Using the same set of data, the WPA Slave Narratives Project, connections were explored using Voyant 2.0, kelper.gl, and Palladio. Each of the tools offered a way of understanding the information and had value as a stand alone analysis, however they also complement each other.

The tools revealed different aspects of the narratives. After working with Palladio and seeing the connections between topics and demographics, I think taking the topics that appeared most often (or least) in the Alabama example would be interesting to run the data again in Voyant 2.0 and see what the context was. I was surprised at the religion-Sunday-church-wedding -baptism connections, and think bringing Voyant’s analysis in for similar trends in other states may offer insight. This struck me as baptism was the item mentioned least, and may be worth exploring, as both religion and church were fairly strong topics, why was baptism less so given it is a tenet for belonging to many churches. It may also suggest church as a safe space or as a monitored space, threads worth pulling using text analysis combined with network analysis. 

After listening to a graduate student presenting her research on African Americans and migration in South Carolina this afternoon, I think I would like to test use of the three tools in conjunction with each other on a related question. Specifically, I would like to use them to develop visualizations of the movement of the formerly enslaved, whether they traveled a few minutes from the place of enslavement or chose to build a life in another location completely. The Alabama sample, for example, shows limited movement, until you take into consideration that the interviews took place over a short period, in Alabama. Kepler.gl and Palladio can show where those that remained in Alabama did not tend to travel far, but the entire data set would help expand the geospatial view. Voyant would be a useful tool to add the context to the data, and potentially add some information on those who left as part of the Great Migration.

Palladio and Networks

Palladio is open source software that allows users to visualize data through mapping, graphs, and tables. Data can be represented as points on a map, connected by arcs and scalable, in a list or table form and in a gallery grip layout for organization. There are also timeline and timespan filters that it did not explore.

One of the interesting features in graphs is the dynamic interface, allowing manipulation of the network graph to isolate or drill down on nodes and connections. For example, in the graph below, the topics and work performed data sets were compared highlighting some of the differences in experience and points of view (Figure 1)

Figure 2 shows the relationship between where the person was interviewed compared to where they were enslaved. The grouping suggests limited movement between the two locations, and highlights an area to compare the WPA interview data to census data to understand the paths of movement and whether the Alabama data set is representative of all the movement post-liberation. Comparing this to census data from 1860/1880 and 1930 might be an interesting follow up to this visualization.

Palladio was fairly straightforward to use, with an excellent tutorial page to help with adding data, data structure and the different functions of the software.

Mapping with Kepler.gl

In the digital humanities, mapping offers another way to provide information in a way that engages the audience. Histories of the National Mall, for example, uses a map interface as a gateway for visitors to explore the Mall in a different way. As the developer explained, there are no markers about the history of the space – the Omeka based site serves as digital historical markers.  

Digital mapping tools are powerful for visualizing data. Through digital mapping, relationships, change over time, aggregation or dispersion, and paths/vectors can be highlighted or explored. In an exercise with the Kepler.gl, layers and filters were added to datasets to understand patterns and identify possible relationships between the data (interview information from the WPA Slave narratives).  Some of the questions it left me with were regarding the clusters of interviews geographically, I am thinking I should have played with adding towns and cities, to look closer at what communities may have been targeted by the interviews, or whether it was a random distribution of the survivors of enslavement.

The timeline map also raised some interesting questions as the interview dates clustered in the spring and summer, vice an even distribution across the dates. It leads to the question of were the interviews rushed, or funding running out, or was it just the time available (See map below).

Text Mining with Voyant

I understood text mining as a concept, using Voyant, with all its highs and lows, helped my comprehension. I realized my knowledge was thin, familiar with word clouds and graphing, working with Voyant, and I am curious about other tools as well now, I was able to delve into the context, and found some surprises.

Sinclair and Lockwood discussed the two basic questions for text mining:

  • a means of taking linguistic and semantic characteristics and seeing the different uses and context
  • taking unfamiliar work and making it understandable

The experience I had exploring the Kentucky dataset supports the idea. For Kentucky, I ended up with this cirrus

 

The prevalence of the word “War” surprised me, until I dug into the context. War was discussed as part of the story of participants, in the context of the Civil War, and as a verb. War, it turned out, also meant was. Simply looking at the word cloud or even the trends, would not have shown the other meaning, and just left it as the interviewees must have been deeply invested in the war.

 

Voyant was a little challenging sometimes, with functions working some times and not others, but the value of the tools outweighed any issues with the program.

The text mining exercise was interesting and made me think more about what can be learned, and how little I knew at the start.

Why Metadata Matters

Metadata, in simplest terms is the data about data. Metadata provides a description of a Web resource by assigning attributes to the object. www.dublincore.org uses the example of how libraries use metadata as part of the catalog: title, author, date of creation or publication, etc.  The site goes on to explain that part of the stressful information overload frequently experienced during a search is undifferentiated digital data. 

A Web resource’s basic information follows the basic idea of the library catalog, expanding it out to encompass more information. The Dublin Core Standards offer a framework for elements using terms that were developed across multiple user communities as a means of simplifying metadata.

Metadata attached to a photograph of a coffeemaker adds information for perspective. By offering information on the physical dimensions, it helps visualized the item outside of the flattened digital format.

Database Review

Everyday Life & Women in America

Home – Everyday Life & Women in America (gmu.edu)

Everyday Life & Women in America provides access to primary source material from the Sallie Bingham Center for Women’s History and Culture at  Duke University and The New York Public Library. The collection contains periodicals, pamphlets, monographs and broadsides spanning the period of 1800-1920, which can be viewed and searched multiple ways. The collection is curated to provide “a thematically unified but eclectic range of material,” a seemingly contradictory goal. Searching the database is a little trickier than advertised, as the site repeats that there is an option to view the documents thematically, but it is extremely difficult to determine how to do so. The basic search allows filtering by date and date range,document type and library or archive. Search results offer both a list of documents and a tab for secondary resources, essays commissioned to offer more context to some of the material. Advanced search is available using Boolean search parameters. 

The search guide offered does a thorough job of walking a user through how to search as well as some different ways to think about language in the search. “Selection and Language” offers the warning that due to the time period, language that is offensive today will frequently be used, but may also be necessary and will return search results that may not be found otherwise. They frame the challenges as it is “problematic” but may be critical to finding hidden narratives (there is a section in the searching guide discussing some ways to work around this as well). The search guide also has an excellent section “Find People” that guides the user through language and filter options to help look for people that do not often stand out in the collection. The searching guide also directs the user to three ways to browse the collection: View Documents, a list of documents; Search Directories, which leads to a search by Library of Congress Subject Heading from a drop down; and Thematic Areas. The Thematic Areas are disappointing, as the site describes them as guides to highlight major themes, where in reality they are paragraphs with a few examples. They do come with the disclaimer that documents are not tagged that way, it seems like a lost opportunity.

The collection is 700 plus items published between 1800 and 1920 in the United States. The publisher is Adam Matthew Digital Limited, published first in 2007 with updates to the platform and content in 2017 and 2023 (https://www-everydaylife-amdigital-co-uk.mutex.gmu.edu/introduction/publication-details). Images can be downloaded as the full document in PDF, Current transcript (if available), and the content metadata. Each item is full text searchable. 

Much of the collection comes from the Sallie Bingham Center for Women’s History and Culture at Duke University whose mission is the acquisition and preservation of materials about the lives of women. The Center started with the endowment of an archivist and has expanded to a permanently endowed center in 1993. As a reviewer noted, some of the database skews toward documents from the South, however the partnership with the New York Public Library offers some balance as well as the Town Topics periodical on Gilded Age society in New York. The documents do not list sources of digitization, nor do either of the sources specify how they were digitized.

The database was reviewed twice, one in 2008 and in 2013. Bothe reviews rated the database as useful and worth using. Both commented on the Chronology feature, color coded by category for context, and while it is nifty visually, it is of limited use as the entries do not link back to resources.

  • http://mutex.gmu.edu/login?url=https://www.proquest.com/scholarly-journals/everyday-life-women-america-c-1800-1920/docview/1325052624/se-2?accountid=14541
  • http://mutex.gmu.edu/login?url=https://www.proquest.com/trade-journals/everyday-life-women-america/docview/225720643/se-2?accountid=14541

The database is available by subscription through libraries or educational institutions. The homepage will load without subscription, however only the introduction page is accessible, navigation to any other page requires a login. Copyright information is contained in the metadata for the item, there is no guidance on citations.

Overall, it is a rich database that offers different perspectives on the period. At under 800 items, it is definitely curated, but with an eye toward introducing the depth writing about women’s lives in the era.

Guide to Digitization

In Art History, many years ago, I was shown Dali’s Persistence of Memory. It covered the screen in the auditorium, and I have seen it on posters and t-shirts for years since.  It was the ‘80s, so the image was a projection of a slide from a photograph, but one elements applies to thinking about digitization.  The painting, it turns out, is not a large grandiose work. At the Museum of Modern Art in New York, a viewer will discover Dali’s work is 9.5 inches by 13 inches, making it only a little larger than a sheet of paper.

Salvador Dalí. The Persistence of Memory. 1931 | MoMA

The virtual and reality of the Dali painting demonstrates some of the thought processes required when determining what to digitize and how.   Digitization allows the painting to be seen and studied regardless of location, but it both enhances (allows closer looks at the elements of the piece) and detracts (distorts the size and detail of the physical work) from the object itself.  Digitizing an object is a balance between which properties are transmittable and which are not.   Scale and texture are harder to capture digitally. Most of all, context and bias are impossible to capture. When culturally sensitive objects are being considered for digitization, the potential harm and the bias of the organization digitizing the object have to be weighed against the value of digitizing an object and the purpose of making a digital image.

Conway’s discussion of manipulation of images for different effects was an interesting discussion of the value of an exact copy and what can be done to make a digital image either more useful or more faithful to the original creator.  For example, the tone adjustment case he used allowed the viewer to see into the home and understand more about the family and its living conditions than the original print.

Video and 3D imaging allows digitization with more perspective. Combining 3D imaging with 3D printing, museums can offer visitors with sensory or vision issues and option to have 3D models of objects to touch and understanding.

The ability to digitize objects allows a wider variety of uses than an object solely in situ. It has the potential to widen the reach of an object, as well as manipulation to develop a better understanding of the object its place in time and space.

Sources for finding usable Digital data 6 of 6

Library of Congress

https://www.loc.gov/

Using Items from the Library’s Website: Understanding Copyright  |  Legal  |  Library of Congress (loc.gov)

The Library of Congress(LOC) contains millions of books, photographs, papers, maps, films, sundry printed materials and recordings. The digital collection covers a large portion of these. The digital collection is predominantly in the public domain or no known copyright restrictions, items the LOC has permissions for or those with special access. The rights category that each image falls into is clearly listed in the item information and metadata.