Showing posts with label storage architecture. Show all posts
Showing posts with label storage architecture. Show all posts

Friday, July 20, 2012

Amazon's loss is SDSC's gain

One of the recent Amazon Web Services (AWS) power outages has left some of my EBS volumes in an inconsistent state.  If these were simple volumes, each containing a filesystem, then the fix is easy:  just dismount the filesystem, run fsck to check it, and then remount the filesystem after it has been fixed.  We have done this on several of our EC2 instances that had inconsistent volumes.

Unfortunately, for these particular volumes we have bonded them together to form a virtual RAID.  And this RAID is used as a single multi-TB filesystem which is much bigger than fsck can handle.  So we are kind of stuck.

One option would be to newfs the big filesystem, and to move the several TBs of content back into AWS, but that would be very slow.  And if there is another power outage......

So instead we called up our pals at Duracloud and asked them if they could help us enable replication of our content to a second provider.  (The first provider is - ironically - AWS.  But their S3 service, not their EC2/EBS service.)  They said they'd be happy to help, and, in fact, they will starting to replicate our content later this same week.  (Now that's service!)

The new copy of our content will now be replicated in...... SDSC's storage cloud.  This really brings us full circle at ICPSR since our very first off-site archival copy was stored at SDSC. Back then (like in 2008) it was stored in their Storage Resource Broker (SRB) system, and we used a set of command-line utilities to sync content between ICPSR and SDSC.  

The SRB stuff was kind of clunky for us, especially given our large number of files, our sometimes large files (>2GB), and our sometimes poorly named files (e.g., control characters in file names).  Our content then moved into Chronopolis from SRB, and then at the end of the demonstration project, we asked SDSC to dispose of the copy they had.  But now it is coming back......

Friday, June 22, 2012

AWS power outage aftermath

As it turns out, it doesn't take all that long to run fsck on a large filesystem comprised of multiple AWS Elastic Block Storage (EBS) volumes:


[root@cloudora ~]# df -h /dev/md0
Filesystem            Size  Used Avail Use% Mounted on
/dev/md0              4.9T  2.9T  1.7T  64% /arcstore
 
[root@cloudora ~]# fsck -y /dev/md0
fsck 1.39 (29-May-2006)
e2fsck 1.39 (29-May-2006)
/dev/md0 is mounted.
WARNING!!!  Running e2fsck on a mounted filesystem may cause
SEVERE filesystem damage.
Do you really want to continue (y/n)? yes
/dev/md0 has gone 454 days without being checked, check forced.
Pass 1: Checking inodes, blocks, and sizes
Error allocating icount link information: Memory allocation failed
e2fsck: aborted

saw a few posts about coaxing e2fsck to use the filesystem for scratch space rather than memory, but unfortunately the older version of the program available on this EC2 instance does not support it.

So I think that we may end up blowing away this copy of archival storage and replacing it with a fresh one.

Friday, May 18, 2012

Data Fetish

David Rosenthal has another good post on digital preservation where he quotes a gent who has coined the term "data fetish" for those that think that data should never be thrown away.  Well worth reading, but I'll quote one piece here:
So, we're going to have to throw stuff away. Even if we believe keeping stuff is really cheap, its still too expensive. The bad news is that deciding what to keep and what to throw away isn't free either.
This is so very true.

When my team inherited the responsibility for Archival Storage at ICPSR in 2005-2006, we made a choice to throw away a ton of stuff.  Paper hardcopies of electronic files.  Magnetic tapes that had been read successfully  to spinning disk.  Old office supplies.  Leases to storage lockers. .... After the smoke cleared we had about 5TB of content moved from a large array of tape types to spinning disk.  And this formed the core of ICPSR's Archival Storage.

I think a key question for every archive to answer is "Why are we keeping this?"

If the answer is "We're not sure" (even with some oblique wording), then it is time for that item to be discarded.

Wednesday, November 16, 2011

DuraCloud Archiving and Preservation Webinar

Shameless self-promotion alert...

The nice folks at DuraSpace have published the audio and video from the recent webinar that Michele Kimpton (CEO, DuraSpace) and I gave on DuraCloud.


Michele spends the first 5-10 minutes talking about the business case behind DuraCloud, and then I spend about 30 minutes talking about ICPSR and how we came to use DuraCloud to store a copy of our archival holdings.

Wednesday, October 26, 2011

Using DuraCloud for Archiving and Preservation

I'll be joining Michele Kimpton, CEO of DuraSpace, on a webinar next Wednesday (November 2, 2011).  Our topic is DuraCloud, and how one can use this cloud-based service as part of one's digital preservation strategy.

I think sometimes people will view the cloud as an alternative to keeping and maintaining local copies, but at ICPSR we're using the cloud as an easy-to-manage storage location to supplement more conventional locations, such as local NAS storage and the University of Michigan's "Value Storage" service.

Here is a copy of the invite that went out via email:


DuraSpace

You are invited to attend the following event:

Using DuraCloud for Archiving & Preservation
Wednesday, November 2, 2011
1:00p.m. - 2:00p.m. Eastern Standard Time

Presented By:
Michele Kimpton, DuraSpace Chief Executive Officer & DuraCloud Project Director &
Bryan Beecher, Director of Computing & Network Services,Interuniversity Consortium for Political and Social Research (ICPSR)


Having a hard time keeping up with current preservation and archiving practices?
Are you finding the task of archiving your content complicated, costly and confusing?
Then you need to join us for a free webinar that details how DuraCloud can be part of your preservation and archiving solution.

This webinar will discuss how to use DuraCloud as a component of your archiving and preservation strategy. An overview of the service will include what it is, how it works, and the benefits it has to offer. Additionally, Bryan Beecher, Director of Computing & Network Services at ICPSR, will present ICPSR's preservation and archivingstrategy. Bryan will share how DuraCloud and other methods have been implemented to meet ICPSR's preservation and archiving goals.

If you are interested in attending please thoroughly complete the registration process (below) to receive your unique login url. Be sure to SAVE the return email you receive from Infinite Conferencing as it will include your unique login information.
 

Below is the call-in information for the event.
Via Skype (Free, World): Dial +9900827047086940 
Via Phone (Toll, US): Dial +1(201)793-9022 Enter Room Number: 7086940


The maximum capacity for this web seminar is 99 participants. The event will be recorded and slides will be available for viewing after the event at http://duraspace.org/web_seminars
.

Please contact Kristi Searle at ksearle@duraspace.org
 with any questions.

**Please be aware of Infinite Conferencing System Requirements:
-          Internet connection speed of 128 kbps or higher is recommended
-          Microsoft Windows XP, Vista, Windows 7, or Server 2003
-          Internet Explorer 6.0 SP2, 7.0, 8.0 & 9.0, Firefox 3.0x/3.5, 4, 5 and Chrome 12 browsers
-          Apple Mac with Intel CPU, Mac OS X 10.5/10.6, Safari 4.x, 5.x or Firefox 3.x, Java 1.5+
-          Linux, Unix, or Solaris with Mozilla 1.0+
-          Cookies and Scripting enabled in browser

Friday, September 30, 2011

Designing Storage Architectures for Digital Preservation

This is the second part of a two-part post about the 2011 Designing Storage Architectures for Digital Preservation meeting hosted by the Library of Congress.  The first part can be found in this post.


The second day began with a second session on Power-aware Storage Technologies.

Tim Murphy (SeaMicro) spoke about his lower-power server offering, noting that "Google spends 30% of its operating expenses on power" and how it "costs more to power equipment than to buy it."  Dale Wickizer (NetApp) gave a talk on how Big Data is now driving enterprise IT rather than enterprise applications or decision support.  Ethan Miller (Pure Storage) described his lower-power, flash-based storage hardware, and how a combination of de-dupe and compression makes it cost comparably to enterprise hard disk drives (HDD).  Dave Anderson (Seagate) spoke about HDD security and how new technology aimed at encryption may make sense for digital preservation applications too.

The theme of the next session was New Innovative Storage Technologies.

David Schissel (General Atomics) presented an overview of their enhanced version of the old Storage Resource Broker (SRB) technology which they call Nirvana.  Bob [did not catch his full name] (Nimbus Data) described his flash-based storage array, and how it applied the same techniques as conventional disk-based storage arrays, but with flash instead.  John Topp (Splunk) described his product which struck me as a giant indexer and aggregator of log file content.  Sam Thompson (IBM) spoke about BigSheets, which layers a spreadsheet metaphor on top of technologies like nutch, mapreduce, etc.

This theme continued into the next session.

Chad Thibodeau (Cleversafe) described his technology for authenticating to cloud storage in a more secure manner by distributing credentials across a series of systems.  Jacob Farmer (Cambridge Computer) proposed adding middleware between content management systems and raw storage to make it easier to manage and migrate content.  R B Hooks (Oracle) presented an overview of trends in storage technology, and noted that the consumer market, not IT, will drive flash technology.  Marshall Presser (Greenplum) spoke about I/O considerations in data analytics.

The day ended with two closing talks.

Ethan Miller (UC Santa Cruz this time) spoke about the need to conduct research into how archival storage systems are actually used.  He described results from a pair of initial studies.  In the first, access was nearly non-existent, except for a one-day period where Google crawled the storage, and this one day accounted for 70% of the access during the entire time period of the study.   In the second, 99% of the access was fixity checking.  [I think this is how ICPSR archival storage would look.] 

David Rosenthal (LOCKSS Project, Stanford University) presented a still evolving model of how one computes the long-term storage costs of digital preservation.  The idea is that this model could be used to answer questions about whether to buy or rent (cloud), when to upgrade technologies, and so on.  You can find the full description of the model at David's blog here.



Wednesday, September 28, 2011

Designing Storage Architectures for Digital Preservation

I attended the 2011 edition of the Library of Congress' Designing Storage Architectures for Digital Preservation meeting (link to 2010 meeting).  Like previous events, this meeting was scheduled over two days, and featured attendees and speakers from industry, higher education, and the US government.  This post will summarize the first day of the meeting, and I'll post a summary of the second day later this week.

The meeting was held in the ballroom of The Fairfax on Embassy Row in Washington, DC.  About 100 people attended the event which began at noon on Monday, September 26, 2011.  As at past meetings the first hour was devoted to registration and a buffet lunch.

The program began at 1:00pm with a welcome from Martha Anderson, who leads the National Digital Information Infrastructure and Preservation Program (NDIIPP) program for the Library of Congress (LC).  She noted that since its inception in 2000, the program has funded 70 projects spread across 200 organizations, and that it is valuable for people to be able to step out of the office for a short time to step back, see the big picture in digital preservation, and get fresh perspectives.  She described how change is the driving force in digital preservation, and characterized one big change as a shift from indexing content to processing content.

Two "stage setting" presentations followed.

Carl Watts (LC) described a massive migration underway at the LC where 500TB of content was moving from one storage platform to another.  Henry Newman (Instrumental) described challenges facing the digital preservation community:  data growth is greatly exceeding growth in hardware speed and capacity; POSIX has not changed in many years; nomenclature is not used consistently between digital preservation practitioners and vendors; and, the total cost of ownership for digital preservation is not well understood.

The theme of the first session was Case Studies from Community Storage Users and Providers.

Scott Rife (LC) described the video processing routine used at the LC Packard Campus, which handles over 1m videos and 7m audio files, 7TB/day of content, and 2GB/s of disk access.  Jim Snyder (LC) also spoke about the Packard Campus, noting that he is trying to "engineer for centuries" where one generation hands off to the next generation.  Cory Snavely (University of Michigan) and Trevor Owens (LC) gave an overview of the National Digital Stewardship Alliance (NDSA), and a more detailed report of what has been happening in the Infrastructure Working Group [note - I am a member of that WG] including preliminary results from a survey of members.  Highlights: 87% of respondents intend to keep content indefinitely; 76% anticipate a infrastructure change within three years; 72% want to host content themselves; 50% want to outsource hosting (!); 57% are using or considering use of "the cloud"; and, 60% intend to work through the TRAC process. 

Steve Abrams (California Digital Library) spoke about a "neighborhood watch" metaphor for assessing digital preservation success.  Tab Butler (Major League Baseball Network) updated the audience on the staggering amount of video he manages (2500 hours of HD video each week with multiple copies/versions of many of the hours).  Barbara Taranto (New York Public Library) described a migration where the content doesn't move; only its address changes (in a Fedora repository).  Corey Snavey (Hathitrust this time) gave a second talk, updating the audience on text searching at the Hathitrust.  More memory delivers better performance.  Andrew Woods (DuraSpace) described some of the challenges his team has faced building storage services across disparate cloud storage providers. 

The theme of the next session was Power-aware Storage Technologies.

Hal Woods (HP) forecast a shift to solid state drives (SSD) in the next 2-4 years, and speculated that tape might outlive hard disk drives (HDD).  Bob Fernander (Pivot3) described video as the "new baseline" for content, and warned that we need to stop building Heath Kit style solutions to problems.  Dave Fellinger (DataDirect Networks) advised that the building blocks of digital preservation solutions needed to be bigger, and building with the right-sized block would make it easier to solve problems.  Mark Flournoy (STEC) gave a very nice overview of different SDD market segments, costs, and performance metrics.

Each session included a lengthy question, answer, and comment section, and sometimes lively debate amongst the audience. 

The first day wrapped up a bit after 5pm.

Monday, September 5, 2011

Designing Storage Architectures for Preservation Collections

Tech@ICPSR will be heading to DC in late September to attend another annual installment of Designing Storage Architectures for Preservation Collections hosted by the Library of Congress (link to last year's meeting page).  This promises to be another useful meeting, and as I've done in the past, I'll post some my notes from the meeting.

The past couple of meetings have given considerable attention to the cost of the hardware and software systems that supply the basic storage platform.  There's usually a lot of interesting tidbits in those conversations, but my sense is that the people costs of ingest and curation are the major costs at ICPSR.  I can't tell if that's unusual amongst this crowd (e.g., they have way more content than we do, and it requires far less human touch), or if it is the proverbial elephant in the room that no one mentions.

Tuesday, November 23, 2010

The Cloud and Archival Storage

Price.  Availability.  Services.  Security.

These are the four parameters that I use when deciding where to store one of our archival storage copies.

For me the cloud is just another storage container.  Fundamentally it is no different from a physical storage location except in how it differs across these four dimensions.  In fact, I can conceptualize my "non-cloud" storage locations as storage as a service cloud providers, but where the provider is a lot more local than the big names in "cloud" today:

ICPSR Cloud:  This is the portion of the EMC NS-120 NAS that I use for a local copy of archival storage.  It is very expensive with a reasonably high-level of availability.  It provides very few services; if I want to perform a fixity check of the objects I have stored here, I have to create and schedule that myself.  Because I have physical control over the ICPSR Cloud, I have an irrational belief that it is probably secure, even though I know that ICPSR isn't as physically secure as many other companies at which I have worked.  Certainly ICPSR does not make any statements or guarantees about ISO 27001 compliance.

UMich Cloud:  This is a multi-TB chunk of NFS file storage that I rent from the University of Michigan's Information Technology Services (ITS) organization.  They call it ITS Value Storage.  The price here is excellent, but the level of availability is just a hair lower.  I don't notice the lower level of availability most of the time, but I do perceive it when running long-lived, I/O-intensive applications.  Like my own cloud, this one has no services unless I deploy them myself.  Because I do not have physical control over the equipment, or even know exactly where the equipment is (beyond a given data center), it feels like there is less control.  ITS makes no promises about ISO 27001 compliance (or promises about other standards), but my sense is that their controls and physical security and IT management processes must be at least as good as mine.  After all, they are managing many, many TBs for many different university departments and organizations, including themselves.

Amazon Cloud:  This is a multi-TB chunk of Elastic Block Storage (EBS) that I rent from Amazon Web Services.  I use EBS rather than the Simple Storage Service (S3) because I want the semantics of a filesystem so that I don't have to worry about things like files that are large or that have funny characters in their names.  The price here is good, better than my EMC NAS, but not as good as the ITS Value Storage.  The availability is quite good overall, but, of course, the network throughput between ICPSR and AWS is nowhere near as good as intra-campus networking, and it is even worse for the AWS EU location.  The services are no better and no worse than my own cloud or the UMich cloud.  Like the ITS Value Storage service I have no control over the physical systems, and I know even less about their physical location.  Amazon says that it passed a SAS 70 audit, and recently received an ISO 27001 certification.  This seems to be a better security story than anyone else so far.

DuraCloud:  Unlike the other clouds, I'm not using this one for archival storage; it is still in a pilot phase.  The availability is similar to plain old AWS (which hosts the main DuraCloud service), and the price is still under discussion.  My expectation is that the level of security is no better (and no worse) than the underlying cloud provider(s), and so depending upon which storage provider one selects, one's mileage may vary.  However, the really interesting thing about DuraCloud is the idea of the services.  If DuraCloud can execute and deliver useful, robust services on top of basic storage clouds, that will be a true value-add, and will make this a very compelling platform for archival storage.

Chronopolis:  Like DuraCloud, this too is not in production yet, and is being groomed (I think) as a future for-fee, production service.  I don't have as much visibility here with regard to availability since I am not actively moving content in and out of Chronopolis; most of the action seems to be taking please under the hood between the storage partner locations.  My sense is that the level of security is probably similar to the UMich Cloud since the lead organization, the San Diego Supercomputer Center, is in the world of higher education, like UMich, but it may well be the case that they have a stronger security story to tell, and I just don't know it.  And like DuraCloud, my sense is that it will come down to services:  If Chronopolis builds great services that facilitate archival storage, that will make it an interesting choice.

Wednesday, October 13, 2010

Designing Storage Architectures for Digital Preservation - Day Two, Part Two

The final session of the conference featured six speakers.

  1. Jimmy Lin (University of Maryland) is spending some time at Twitter, and described their technology stack: hardware, HDFS, Hadoop, and pig, which he described as the "perl/python of big data."
  2. Mike Smorul (University of Maryland) gave an overview of their "time machine for the web" and the challenges of managing a web archive
  3. John Johnson (Pacific Northwest National Laboratory) proposed that the scientific process has changed in that data produced by computation is now one of the drivers for creating and testing new theories
  4. Leslie Johnston (Library of Congress) spoke briefly about an IBM emerging technology called "big sheets"
  5. Dave Fellinger (DataDirect Networks) urged the audience to "don't be afraid to count machine cycles" when analyzing storage systems for bottlenecks that increase service latency
  6. Kevin Kambach (Oracle) finished the session with industry notes about large data
The day then concluded with two final talks.  One was from Subodh Kulkarni (Imation) who gave an overview of storage technology from magnetic tape to hard disk, and the other was from David Rosenthal (LOCKSS) who gave an abbreviated version of his iPres talk, "How Green is Digital Preservation?"  David mentioned a very interesting, large-scale, low-power computing and storage platform being produced by a company called Seamicro.

Thursday, September 9, 2010

Storage Architectures for Digital Collections 2010

I'm heading to DC for the Library of Congress's 2010 edition of their Storage Architectures for Digital Collections event. Should be another interesting meeting, and I'll be sure to include a brief summary after the meeting. (The LoC also produces a very nice, detailed summary, but it usually appears several weeks after the event.)

Wednesday, August 25, 2010

DuraCloud pilot update


Things are starting to move along nicely with ICPSR's participation in the DuraCloud pilot. My early experience with the software and tools is mostly positive, but they are still clearly a bit rough around the edges. For example, I've run into bugs on the login screen of the DuraCloud Admin web app that I would characterize as minor, such as needing to use the Submit button on the screen and not the Enter key on the keyboard for some browsers. That said, the DuraSpace people have been fabulous: It's clear they care a lot about the project and the pilot testers, and they have been very, very responsive.

ICPSR is going to test out three parts of DuraCloud.

One, we'll execute a basic upload test, moving content from ICPSR to a "space" in DuraCloud. For this test I'll be using the DuraCloud Admin tool to create a "space," which is basically the same thing as a "bucket" in Amazon's S3. Then I'll use the DuraCloud "sync tool" to copy a subset of ICPSR's archival content to DuraCloud.

Two, we're going to help spec out a "dashboard" or high-level view that shows the integrity of a collection in DuraCloud, and then execute a "fixity" test to measure performance and reliability. For example, if I have a 1TB collection in DuraSpace, and I have that collection replicated across N cloud storage providers, what's my cost to execute such a test every week?

Three, we're going to test a "replication helper" utility that facilitates replicating content across cloud providers. This is a very compelling service for me. Since we already make extensive use of AWS, using DuraCloud as a front-end to AWS is not very interesting for us; but, if we can use DuraCloud as a single front-end to AWS and RackSpace and Atmos and .... then things get more interesting since it means we don't have to develop expertise with ALL of the cloud providers.

Tuesday, June 15, 2010

DuraCloud Pilot

ICPSR is one of several organizations which are participating in "Round #2" of the DuraCloud pilot. This phase of the pilot begins in September and runs through the end of the calendar year.

I had missed two earlier webinar sessions on the technology and the pilot process, but read through the slides and listened to the audio today. It really looks like DuraCloud will be an interesting project. In brief it is essentially an abstraction layer on top of multiple cloud storage providers, and also implements a series of services, such as a bit-integrity checker.

For our participation I think we'll use our archival content, but nothing that involves any level of confidentiality. So fair game would be previous (and current) versions of public-use datasets and documentation files (codebooks), and perhaps snapshots of study-level metadata encapsulated in DDI XML.

If the pilot goes well and DuraCloud looks like an attractive service for making additional preservation copies of materials, a future project might be to design a system for encrypting our more sensitive content so that it too could be placed in cloud storage.

Wednesday, May 19, 2010

Amazon S3 Introduces Reduced Redundancy Storage


Amazon announced a new product today: Essentially they are making their existing Simple Storage Service (S3) product availability with a slightly lower level of availability and durability. The full news is here.

For a place like ICPSR that likes to replicate content to build a robust archival storage solution, this is good news. If the copy I store with Amazon is "only" 99.99% available, I can live that. After all, it's only one copy of many.

We use a similar service from the University of Michigan central IT organization that they brand Value Storage. There are no backups for disaster recovery, but content is replicated between two U-M data centers.

Thursday, April 29, 2010

ICPSR Storage Locations, 2010 edition



I originally posted a diagram of where ICPSR keeps copies of its content back in November 2009. You can find the post here.

The main change since that post is that we've started using a new storage service from the University of Michigan central IT organization, ITS. This is nice because it gives us another off-site storage location, although the geographic diversity from ICPSR is not that good.

It's also likely that we'll be participating in a pilot project to evaluate a cloud-based storage service, and we're also active in a project to build a LOCKSS-based storage network. Both of these projects hold great promise for creating several additional copies of our holdings in secure, durable storage.

Tuesday, September 15, 2009

Library of Congress Meeting on Storage Architectures for Digital Collections


I'm heading to Washington, DC next week to participate in a meeting hosted by the Library of Congress. Laura Campbell, the Associate Librarian for Strategic Initiatives describes it this way:

The meeting will bring together technical industry experts, IT professionals, digital collections and strategic planning staff, government specialists with an interest in preservation, and recognized authorities and practitioners of Digital Preservation. We would like to be able to make progress on identifying the areas that should matter to those who are responsible for the digital content and for those who are responsible for providing the services to manage the content. We are hoping that we can inform their decision-making in the future, and give them confidence and comfort that they are asking the right questions and can understand the answers. We believe these questions and answers are often common to organizations dealing with different types of content used for different purposes, so we expect that the topics will be of broad interest in the community.

This should be an interesting meeting, and is a nice complement to many related activities at ICPSR. For example, earlier this year I participated in Nancy McGovern's Digital Preservation Management workshop, which was a wonderful overview of the latest news and best practices, and helped fill in several gaps in my understanding of the OAIS reference model. (Nancy is also my colleague at ICPSR and serves as our Digital Preservation Officer.)