DSpace Repository

The Cost of a WARC: Analyzing Web Archives in the Cloud

The Cost of a WARC: Analyzing Web Archives in the Cloud

Show full item record

Title: The Cost of a WARC: Analyzing Web Archives in the Cloud
Author: Deschamps, Ryan
Fritz, Samantha
Lin, Jimmy
Milligan, Ian
Ruest, Nick
Abstract: The value of web archives to support scholarship in the humanities and social sciences is slowly being realized by the increasing availability of scalable tools and platforms. The cost of providing scholarly access is a critical component of developing a long-term sustainability strategy. This paper attempts to answer a straightforward question:\ How much does it cost to analyze web archives in the cloud? To make this question more concrete, we examine the creation of three derivatives (extraction of collection statistics, full text, and the webgraph) that serve as the starting points of many scholarly inquiries. Our analysis shows that these typical derivatives costs around US$7 per TB using our Archives Unleashed Toolkit. We describe in detail the methodology and assumptions made to arrive at this figure. To our knowledge, we are the first to quantify the economics of scholarly access to web archives, and we believe that this information is valuable for service planning by archives, libraries, and other institutions.
Sponsor: This work was primarily supported by the Andrew W. Mellon Foundation, with additional funding from Start Smart Labs, the Natural Sciences and Engineering Research Council of Canada, the Social Sciences and Humanities Research Council of Canada, and the Ontario Ministry of Research and Innovation's Early Researcher Award program. We'd like to thank our content partners and Raymie Stata for comments on an earlier draft.
Subject: web archives
computational analysis
cloud computing
Type: Article
URI: http://hdl.handle.net/10315/36158
Citation: Ryan Deschamps, Samantha Fritz, Jimmy Lin, Ian Milligan, and Nick Ruest. “The Cost of a WARC: Analyzing Web Archives in the Cloud.” Proceedings of the ACM/IEEE Joint Conference on Digital Libraries, Vol. 19 (2019).
Date: 2019

Files in this item

This item appears in the following Collection(s)