Showing posts with label conferences. Show all posts
Showing posts with label conferences. Show all posts

Friday, September 30, 2011

Designing Storage Architectures for Digital Preservation

This is the second part of a two-part post about the 2011 Designing Storage Architectures for Digital Preservation meeting hosted by the Library of Congress.  The first part can be found in this post.


The second day began with a second session on Power-aware Storage Technologies.

Tim Murphy (SeaMicro) spoke about his lower-power server offering, noting that "Google spends 30% of its operating expenses on power" and how it "costs more to power equipment than to buy it."  Dale Wickizer (NetApp) gave a talk on how Big Data is now driving enterprise IT rather than enterprise applications or decision support.  Ethan Miller (Pure Storage) described his lower-power, flash-based storage hardware, and how a combination of de-dupe and compression makes it cost comparably to enterprise hard disk drives (HDD).  Dave Anderson (Seagate) spoke about HDD security and how new technology aimed at encryption may make sense for digital preservation applications too.

The theme of the next session was New Innovative Storage Technologies.

David Schissel (General Atomics) presented an overview of their enhanced version of the old Storage Resource Broker (SRB) technology which they call Nirvana.  Bob [did not catch his full name] (Nimbus Data) described his flash-based storage array, and how it applied the same techniques as conventional disk-based storage arrays, but with flash instead.  John Topp (Splunk) described his product which struck me as a giant indexer and aggregator of log file content.  Sam Thompson (IBM) spoke about BigSheets, which layers a spreadsheet metaphor on top of technologies like nutch, mapreduce, etc.

This theme continued into the next session.

Chad Thibodeau (Cleversafe) described his technology for authenticating to cloud storage in a more secure manner by distributing credentials across a series of systems.  Jacob Farmer (Cambridge Computer) proposed adding middleware between content management systems and raw storage to make it easier to manage and migrate content.  R B Hooks (Oracle) presented an overview of trends in storage technology, and noted that the consumer market, not IT, will drive flash technology.  Marshall Presser (Greenplum) spoke about I/O considerations in data analytics.

The day ended with two closing talks.

Ethan Miller (UC Santa Cruz this time) spoke about the need to conduct research into how archival storage systems are actually used.  He described results from a pair of initial studies.  In the first, access was nearly non-existent, except for a one-day period where Google crawled the storage, and this one day accounted for 70% of the access during the entire time period of the study.   In the second, 99% of the access was fixity checking.  [I think this is how ICPSR archival storage would look.] 

David Rosenthal (LOCKSS Project, Stanford University) presented a still evolving model of how one computes the long-term storage costs of digital preservation.  The idea is that this model could be used to answer questions about whether to buy or rent (cloud), when to upgrade technologies, and so on.  You can find the full description of the model at David's blog here.



Wednesday, September 28, 2011

Designing Storage Architectures for Digital Preservation

I attended the 2011 edition of the Library of Congress' Designing Storage Architectures for Digital Preservation meeting (link to 2010 meeting).  Like previous events, this meeting was scheduled over two days, and featured attendees and speakers from industry, higher education, and the US government.  This post will summarize the first day of the meeting, and I'll post a summary of the second day later this week.

The meeting was held in the ballroom of The Fairfax on Embassy Row in Washington, DC.  About 100 people attended the event which began at noon on Monday, September 26, 2011.  As at past meetings the first hour was devoted to registration and a buffet lunch.

The program began at 1:00pm with a welcome from Martha Anderson, who leads the National Digital Information Infrastructure and Preservation Program (NDIIPP) program for the Library of Congress (LC).  She noted that since its inception in 2000, the program has funded 70 projects spread across 200 organizations, and that it is valuable for people to be able to step out of the office for a short time to step back, see the big picture in digital preservation, and get fresh perspectives.  She described how change is the driving force in digital preservation, and characterized one big change as a shift from indexing content to processing content.

Two "stage setting" presentations followed.

Carl Watts (LC) described a massive migration underway at the LC where 500TB of content was moving from one storage platform to another.  Henry Newman (Instrumental) described challenges facing the digital preservation community:  data growth is greatly exceeding growth in hardware speed and capacity; POSIX has not changed in many years; nomenclature is not used consistently between digital preservation practitioners and vendors; and, the total cost of ownership for digital preservation is not well understood.

The theme of the first session was Case Studies from Community Storage Users and Providers.

Scott Rife (LC) described the video processing routine used at the LC Packard Campus, which handles over 1m videos and 7m audio files, 7TB/day of content, and 2GB/s of disk access.  Jim Snyder (LC) also spoke about the Packard Campus, noting that he is trying to "engineer for centuries" where one generation hands off to the next generation.  Cory Snavely (University of Michigan) and Trevor Owens (LC) gave an overview of the National Digital Stewardship Alliance (NDSA), and a more detailed report of what has been happening in the Infrastructure Working Group [note - I am a member of that WG] including preliminary results from a survey of members.  Highlights: 87% of respondents intend to keep content indefinitely; 76% anticipate a infrastructure change within three years; 72% want to host content themselves; 50% want to outsource hosting (!); 57% are using or considering use of "the cloud"; and, 60% intend to work through the TRAC process. 

Steve Abrams (California Digital Library) spoke about a "neighborhood watch" metaphor for assessing digital preservation success.  Tab Butler (Major League Baseball Network) updated the audience on the staggering amount of video he manages (2500 hours of HD video each week with multiple copies/versions of many of the hours).  Barbara Taranto (New York Public Library) described a migration where the content doesn't move; only its address changes (in a Fedora repository).  Corey Snavey (Hathitrust this time) gave a second talk, updating the audience on text searching at the Hathitrust.  More memory delivers better performance.  Andrew Woods (DuraSpace) described some of the challenges his team has faced building storage services across disparate cloud storage providers. 

The theme of the next session was Power-aware Storage Technologies.

Hal Woods (HP) forecast a shift to solid state drives (SSD) in the next 2-4 years, and speculated that tape might outlive hard disk drives (HDD).  Bob Fernander (Pivot3) described video as the "new baseline" for content, and warned that we need to stop building Heath Kit style solutions to problems.  Dave Fellinger (DataDirect Networks) advised that the building blocks of digital preservation solutions needed to be bigger, and building with the right-sized block would make it easier to solve problems.  Mark Flournoy (STEC) gave a very nice overview of different SDD market segments, costs, and performance metrics.

Each session included a lengthy question, answer, and comment section, and sometimes lively debate amongst the audience. 

The first day wrapped up a bit after 5pm.

Monday, September 12, 2011

DuraCloud Pilot talk at Educause 2011

Our colleague from DuraSpace, CEO Michele Kimpton, is giving a talk at this year's Educause conference.  Here is a link to the abstract.

The topic is the DuraCloud pilot that concluded earlier this year, and in which ICPSR was an active participant.  I've been impressed with the level of commitment and the service orientation of the entire DuraCloud team, and we've since moved from "pilot user" to "customer."


Wednesday, June 1, 2011

ICPSR NSF INTEROP project heads to Vancouver for IASSIST 2011

Here is a sneak preview at the poster we've put together for the IASSIST 2011 conference for our NSF INTEROP EAGER project investigating Fedora as a repository solution for social science datasets.

Mary Vardigan, Assistant Director of ICPSR, and Director of Collection Delivery, will be presenting the poster at the IASSIST 2011 conference's poster session on Thursday, June 2, 2011.

In the poster we've tried to capture the thinking that went into our object model for social science research data and documentation, and how we mapped that to Fedora's Content Model Architecture.  We also highlight the other two deliverables of the project, a tool for generating FOXML-format objects for ingest into Fedora, and a suite of still-evolving services that may be applied to the objects.

Wednesday, October 13, 2010

Designing Storage Architectures for Digital Preservation - Day Two, Part Two

The final session of the conference featured six speakers.

  1. Jimmy Lin (University of Maryland) is spending some time at Twitter, and described their technology stack: hardware, HDFS, Hadoop, and pig, which he described as the "perl/python of big data."
  2. Mike Smorul (University of Maryland) gave an overview of their "time machine for the web" and the challenges of managing a web archive
  3. John Johnson (Pacific Northwest National Laboratory) proposed that the scientific process has changed in that data produced by computation is now one of the drivers for creating and testing new theories
  4. Leslie Johnston (Library of Congress) spoke briefly about an IBM emerging technology called "big sheets"
  5. Dave Fellinger (DataDirect Networks) urged the audience to "don't be afraid to count machine cycles" when analyzing storage systems for bottlenecks that increase service latency
  6. Kevin Kambach (Oracle) finished the session with industry notes about large data
The day then concluded with two final talks.  One was from Subodh Kulkarni (Imation) who gave an overview of storage technology from magnetic tape to hard disk, and the other was from David Rosenthal (LOCKSS) who gave an abbreviated version of his iPres talk, "How Green is Digital Preservation?"  David mentioned a very interesting, large-scale, low-power computing and storage platform being produced by a company called Seamicro.

Tuesday, October 12, 2010

Designing Storage Architectures for Digital Preservation - Day Two, Part One

The Library of Congress hosted a two-day meeting on September 27 and 28, 2010 to talk about technologies, strategies, and techniques for managing storage.  Like the 2009 meeting, which I also attended, the meeting is heavily focused on IT and the costs of the the technology.  This was another interesting and valuable meeting, but it always feels like we don't address the elephant in the room:  the cost of all of the people who curate content, create metadata, manage collections, assess content, etc.  This is the report from the second day of the conference.

The morning session of the second day of the conference featured six speakers, many from industry:  
  1. Micah Beck (University of Tennessee - Knoxville) made an argument for "lossy preservation" as a strategy for achieving "good enough" digital preservation in an imperfect world, and suggested that developing techniques for using damaged objects should be part of the archivists' toolkit.
  2. Mike Vamdamme (Fujifilm) gave an overview of their StorageIQ product as a system to augment the reporting and metadata available from conventional tape-based backup and storage systems
  3. Hal Woods (HP) spoke about StorageWorks
  4. Mootaz Elnozahy (IBM) spoke about trends in reliable storage over the next 5-10 years, and predicted that power management requirements will stress hardware causing the rate of MTBF, and the soft error rates of storage to increase.
  5. Dave Anderson (Seagate) also spoke about near-term trends such as a shift to 3TB disks and 2.5" form-factor drives.  He does not see solid state as a factor in the market at this time.
  6. Mike Smorul (University of Maryland) gave a very brief overview of ACE.
The next session featured four more speakers:
  1. Joe Zimm (EMC) was part of Data Domain before being acquired by EMC, and spoke about EMC's block-level de-duplication technology.
  2. Mike Davis (Dell) was part of Ocarina before being acquired by Dell, and spoke about their technology for de-duplication.
  3. Steve Vranyes (Symantec) opined that compression will play a more significant role than de-duplication in easing storage requirements for archives because the use case is very different.
  4. Raghavendra Rao (Cisco) introduced Cisco's network layer de-duplicator.  This seemed like an odd fit in some ways compared to the other products.
Up next - the final post in this series:  the second half of Day Two.

Thursday, September 30, 2010

Designing Storage Architectures for Digital Preservation - Day One, Part Two

The second session of the first day featured technologists from higher education who either operate large archives, or who build systems for operating an archive.

Cory Snavely (University of Michigan, Hathitrust) gave a brief overview of Hathitrust, a repository of digital content shared by many of the Big Ten schools and a few other partners.

Brad McLean (Duraspace) reported on DuraCloud and results from the initial pilot partners.  (ICPSR is part of the current pilot, but was not a member of the original, smaller pilot program.)  He noted theseconcerns about using the cloud for digital preservation:
  1. Some services (such as Amazon's S3) have limits on the size of objects (files)
  2. Bandwidth limits on a per-server basis can impede function and performance
  3. Large files are troublesome
  4. Performance across the cloud can vary widely
  5. (File) naming matters; some storage services limit the type of characters in a name
Brad reiterated a comment made by several others:  A standard for checksums would be good to have.

Matt Schulz (MetaArchive) updated us on the MetaArchive, including a current partnership with Chronopolis.

David Minor (San Diego Supercomputer Center) updated us on the Chronopolis project.  David noted that SDSC is reimplementing its data center, and described three levels of storage in its future architecture:
  1. High-performance storage for scratch content
  2. Traditional filesystem storage
  3. Archival storage
The follow-on discussion included conversations about the right type of interface to access content in archival storage (POSIX, RESTful, object-oriented, etc); the trade-off between using long-lived media and systems for digital preservation v. taking advantage of advances in technology by using short-lived media and systems; and, David Rosenthal reminded everyone that we "... cannot test large systems for zero media failures."

I'll write-up my notes from Day Two early next week.

Wednesday, September 29, 2010

Designing Storage Architectures for Digital Preservation - Day One, Part One

I attended an event on Monday and Tuesday of last week that was hosted by the Library of Congress: Designing Storage Architectures for Digital Preservation.  I also attended the event last year, and so this was my second time attending.

Like last time there were many speakers, each giving a five minute presentation.  Unlike a TED talk where the presentation materials are built specifically to fit well within five minutes, many speakers had conventional slide decks, and raced through them quickly.  Those tended to be the weaker talks since the scope of the material was far too broad for the time allotted.  After a series of presentations there would be group discussion for 15-30 minutes which ran the gamut from interesting and provocative observations to chasing down rabbit holes.

I know the LoC will post complete information about the event, but here is my abbreviated version.  I've tried to hit what I considered to be the highlights, and so the reader should know that this report isn't complete.

The session opened with a video that argued that the Internet gives us more opportunity to innovate since it lowers the barrier for one's "hunches" to "collide" with those of another, and that innovation occurs when two or more good ideas come together.  Henry Newman then gave a framing overview for the meeting that included these interesting points.

  1. IT components are changing/improving at different rates; for example, processors are getting faster more quickly than buses are getting faster
  2. The preservation community and the IT community use different language to talk about archival storage
  3. Preservation TCO is not well understood
  4. The consumer market is driving the storage industry, not the enterprise market
The first of two sessions featured "heavy users" who spoke about some the challenges they faced.  The speakers included Ian Soborhoff (NIST), Mark Phillips (University of North Texas), Andy Maltz (Academy of Motion Picture Arts and Sciences), Ethan Miller (University of California - Santa Cruz), Arcot "Raja" Rajasekar (San Diego Supercomputer Center), Tom Garnett (Biodiversity Heritage Library), Barbara Taranto (New York Public Library), Martin Kalfatovic (Smithsonian Institution), and Tab Butler (Major League Baseball Network).  Highlights of their presentations and the follow-on discussion:
  • Experienced recent sea change where it was no longer possible to forecast storage needs whatsoever
  • "Archival storage... whatever that is."
  • Pergamum tome technology looks very interesting for smart, low-power storage
  • iRODS main components:  data server cloud, metadata catalog, and the rule engine
  • "Open access is a form of preservation."
  • If one needs N amount of space for one copy of archival storage, one also needs 2 x N or 3 x N for the ingest process
  • The "long now"
  • The MLB Network data archive will consume 9000 LTO-4 tapes for storage in 2010.
  • "Digital preservation sounds like hoarding."
  • "After our content was indexed by Google, usage went up 10x."
  • Data recovery from corrupted files is a digital preservation concern.
  • Forensics of a format migration is an effective tool for finding problems in a repository.
Next:  the second session of Day One.

Thursday, September 9, 2010

Storage Architectures for Digital Collections 2010

I'm heading to DC for the Library of Congress's 2010 edition of their Storage Architectures for Digital Collections event. Should be another interesting meeting, and I'll be sure to include a brief summary after the meeting. (The LoC also produces a very nice, detailed summary, but it usually appears several weeks after the event.)

Wednesday, October 21, 2009

SUMIT 09 - Annual UMich IT Security Symposium

I attended a very interesting symposium at UMich on Tuesday. It's an annual event called SUMIT, and the focus is on IT-related security. The event includes a series of speakers who have interesting stories to tell, and this year was no exception.

I arrived rather late to the event, and only caught the final part of what appeared to be a very interesting talk by Wade Baker, Verizon Business Security Solutions: Cybercrime: The Actos, Their Actions, and What They're After. Wade's experience has been that data loss is often left undiscovered for five or six months, and often only becomes discovered when that data is used to commit a crime, such as fraud. His sense is that targets are often repositories of information rather than individual systems (e.g., credit companies v. a home PC with information about only a single credit card). He went on to say that most organizations do not know where most of their sensitive data is located; they'll believe that it is located only in areas X and Y, but then discover that someone made a copy in area Z as well. When asked by the audience what single activity is most effective at increasing data security, Wade suggested audits: Organizations often have adequate security policies in place, but all too often they are not followed or enforced, and an audit will reveal this.

The second speaker, Moxie Marlinspike, Institute of Disruptive Technologies, gave a very, very interesting talk entitled Some Tricks for Defeating SSL in Practice. Moxie gave a detailed and clear explanation of a tool he created, sslsniff, and how it can be used in a man-in-the-middle attack to hijack a supposedly secure web connection using SSL. Further, by taking advantage in weak integrity checking by both certificate authorities and certificate-handling software, he demonstrated how one can obtain a "wildcard cert" which allows one to spoof many different web sites. And, as if that isn't scary enough, he also demonstrated how this allows one to inject software onto a machine via automated software-update jobs (e.g., Mozilla's update feature).

The next speaker, Adam Shostack. Microsoft, discussed the economic side of computer security in his talk, New School of Information Security. Adam spoke about how there was a dearth of available data for making decisions about computer security, but that the growing body of "breach data" was improving the situation. Adam pointed to http://datalossdb.org/ as a good example of freely available breach data.

Terry Berg, US Attorney, described the pursuit and resolution of a high-profile case against the spammer, Alan Ralsky, in his talk, To Catch (and Prosecute) a Spammer. In brief, while technology was essential both both perpetrating and later solving the crime, the law enforcement team relied heavily on old-fashioned techniques such as cooperating witnesses to make its case.

The last speaker, Alex Halderman, University of Michigan, discussed a method of defeating secure disk storage through "cold boot" attacks in his talk, Cold-Boot Attacks Against Disk Encryption. It turns out that volatile RAM is not quite so volatile after all, and if one can sufficiently chill a memory chip, one can remove it from a victim PC, install it in a new machine, boot a minimal kernel, and then search the memory for the disk encryption key. Finding the key is easier than one may think because most encryption mechanisms maintain multiple derivatives of the key, which greatly facilitates its theft. The moral of the story is that one should always shutdown a computer or laptop if it contains sensitive data and will be taken through an insecure location (e.g., airport).

Friday, September 25, 2009

IASSIST 2010 Call For Papers

In case you haven't seen this already, here's the call for papers for the next IASSIST meeting:

IASSIST 2010

Social Data and Social Networking:
Connecting Social Science Communities across the Globe
1-4 June 2010
Ithaca, NY, USA

IASSIST 2010, the 36th Annual Conference of the International Association for Social Science Information Service and Technology (IASSIST) will be hosted by the Cornell Institute for Social and Economic Research (CISER) and Cornell University Library (CUL) and will be held at Cornell University, in Ithaca, New York, USA, on 1-4 June 2010.

The theme of this year's conference is Social Data and Social Networking: Connecting Social Science Communities across the Globe. Social science has begun to feel the impact of the dramatic shift in communication patterns globally, where social networking and other digital media trends are changing how social scientists study the world around them. This theme is intended to stimulate discussion about the impact of social networking on the creation, collection, sharing, storage, preservation, dissemination, confidentiality, licensing of, and access to data. Of particular interest is how social connectivity has facilitated multi-site and cross-national social science research.

A webform for submission of proposals will be available on the conference web site: http://ciser.cornell.edu/IASSIST/ from 12 October 2009.


Deadline for submission: 30 November 2009.

Notification of acceptance: 1 February 2010.

For more information about the conference, including travel and accommodation, see the attached PDF Call for Papers or visit the conference web site at:
http://ciser.cornell.edu/IASSIST/ .


IASSIST is an international organization of professionals working in and with information technology and data services to support research and teaching in the social sciences. Typical workplaces include data archives/libraries, statistical agencies, research centers, libraries, academic departments, government departments, and non-profit organizations, see http://www.iassistdata.org for further information.


As usual, I suspect that ICPSR will submit several presentations and papers for the conference. I think an update on either (or both) of the recent tech-oriented grants might be appropriate: using the cloud to deliver sensitive and under-utilized data, and using Fedora to preserve and delivery social science datasets and documentation.