Tech@ICPSR will be giving a talk on cloud computing at the February 1, 2011 LA2M meeting. We'll be talking about the cloud; kind of a high-level overview of what different folks say the cloud is, and some of the consumer- and business-oriented services and systems that live in it.
I'll add a link to the materials shortly after the talk, and, if LA2M adds the video of the talk to their archive, I'll add a link to that as well.
News and commentary about new technology-related projects under development at ICPSR
Showing posts with label icpsr. Show all posts
Showing posts with label icpsr. Show all posts
Monday, January 23, 2012
Monday, December 5, 2011
Collaborators, not depositors
ICPSR should stop accepting deposits.
Instead ICPSR should be recruiting collaborators.
To be sure ICPSR receives a great deal of its content via US Government agencies who have decided to outsource the digital preservation of their content to a trustworthy repository like ICPSR. In this case the relevant contract, grant, or inter-agency agreement makes it clear what content will be coming to ICPSR to be curated and preserved. In some cases the agency has little interest in depositing content ("Isn't that what we pay you for?"), and so the formal act of depositing content falls to the ICPSR staff anyway.
However, we also receive a considerable volume of content through our web portal where the depositor is external. Sometimes we have worked hard to acquire the content, and the deposit is one milestone on a very long road, but other times the content comes to us unsolicited. (I like to call these "drive-by deposits.")
In some cases the depositor is quite eager and able to help ICPSR with much of the curation work: drafting rich descriptive metadata; organizing survey data and documentation into coherent groups; packaging other types of content into logical bundles (such as with our Publication-Related Archive); and, reviewing the data for possible disclosure risks. Depositors may have access to resources like graduate students who can help with these tasks, and if the depositor is also the data producer, then s/he has valuable, unique insight into the data and documentation. Unfortunately ICPSR is not well poised to tap into that expertise and those resources.
What would it take to get there?
ICPSR could separate the transactional step of submitting content (i.e., file upload concurrent with signature) from the iterative step of preparing metadata applicable to the submitted content. In fact, one could even prepare metadata well before the submission transaction if the data producer had the interest and resources to prepare that information, but was not quite ready to share the data yet. And, it would be equally permissible to submit the data for preservation and sharing, and then build the metadata slowly during the weeks and months following the upload.
If the data producer could also export the metadata in machine actionable formats, say, DDI XML for content which maps well to the classic "study" object that ICPSR has curated and preserved for decades, then there may be additional value to the producer. And introducing the structure that comes along with an XML schema like DDI might also be valuable to the producer in terms of thinking about and organizing the documentation, even for his/her own use.
In this world the ICPSR deposit system becomes a much shorter, much simpler web application. And the ICPSR data management infrastructure would need to be opened up -- but with serious access controls -- so that content providers could access, create, and revise their documentation and metadata. But the best thing about this world is that ICPSR gains a lot of collaborators, some who would be quite eager to work with us, I think.
Instead ICPSR should be recruiting collaborators.
To be sure ICPSR receives a great deal of its content via US Government agencies who have decided to outsource the digital preservation of their content to a trustworthy repository like ICPSR. In this case the relevant contract, grant, or inter-agency agreement makes it clear what content will be coming to ICPSR to be curated and preserved. In some cases the agency has little interest in depositing content ("Isn't that what we pay you for?"), and so the formal act of depositing content falls to the ICPSR staff anyway.
However, we also receive a considerable volume of content through our web portal where the depositor is external. Sometimes we have worked hard to acquire the content, and the deposit is one milestone on a very long road, but other times the content comes to us unsolicited. (I like to call these "drive-by deposits.")
In some cases the depositor is quite eager and able to help ICPSR with much of the curation work: drafting rich descriptive metadata; organizing survey data and documentation into coherent groups; packaging other types of content into logical bundles (such as with our Publication-Related Archive); and, reviewing the data for possible disclosure risks. Depositors may have access to resources like graduate students who can help with these tasks, and if the depositor is also the data producer, then s/he has valuable, unique insight into the data and documentation. Unfortunately ICPSR is not well poised to tap into that expertise and those resources.
What would it take to get there?
ICPSR could separate the transactional step of submitting content (i.e., file upload concurrent with signature) from the iterative step of preparing metadata applicable to the submitted content. In fact, one could even prepare metadata well before the submission transaction if the data producer had the interest and resources to prepare that information, but was not quite ready to share the data yet. And, it would be equally permissible to submit the data for preservation and sharing, and then build the metadata slowly during the weeks and months following the upload.
If the data producer could also export the metadata in machine actionable formats, say, DDI XML for content which maps well to the classic "study" object that ICPSR has curated and preserved for decades, then there may be additional value to the producer. And introducing the structure that comes along with an XML schema like DDI might also be valuable to the producer in terms of thinking about and organizing the documentation, even for his/her own use.
In this world the ICPSR deposit system becomes a much shorter, much simpler web application. And the ICPSR data management infrastructure would need to be opened up -- but with serious access controls -- so that content providers could access, create, and revise their documentation and metadata. But the best thing about this world is that ICPSR gains a lot of collaborators, some who would be quite eager to work with us, I think.
Monday, November 7, 2011
October 2011 deposits at ICPSR
Time again for the monthly report of new deposits at ICPSR. Here is the snapshot from October 2011:
The volumes are a bit higher this month, especially the number of files. At least some of the deposits must have been large, containing an unusually large number of files.
In addition to the usual suspects - plain text, stat package formats, MS Word, PDF - we have a very large number of unidentified files this month (2400+ application/octet-stream), and we also have a very small number of interesting formats (images, photoshop).
| # of files | # of deposits | File format |
| 4 | 3 | application/msaccess |
| 46 | 3 | application/msoffice |
| 2496 | 31 | application/msword |
| 2415 | 7 | application/octet-stream |
| 489 | 45 | application/pdf |
| 93 | 16 | application/vnd.ms-excel |
| 6 | 1 | application/vnd.wordperfect |
| 15 | 2 | application/x-dbase |
| 2 | 1 | application/x-empty |
| 1130 | 14 | application/x-sas |
| 1867 | 31 | application/x-spss |
| 1193 | 14 | application/x-stata |
| 1 | 1 | image/gif |
| 1 | 1 | image/jpeg |
| 2 | 2 | image/png |
| 4 | 1 | image/tiff |
| 1 | 1 | image/x-photoshop |
| 10 | 5 | message/rfc8220117bit |
| 151 | 2 | text/html |
| 2 | 2 | text/html; charset=us-ascii |
| 114 | 8 | text/plain; charset=iso-8859-1 |
| 50 | 8 | text/plain; charset=unknown |
| 5047 | 33 | text/plain; charset=us-ascii |
| 67 | 1 | text/plain; charset=utf-8 |
| 4 | 3 | text/rtf |
| 66 | 7 | text/x-c++; charset=us-ascii |
| 1 | 1 | text/x-c++; charset=utf-8 |
| 211 | 7 | text/x-c; charset=us-ascii |
| 98 | 6 | text/xml |
The volumes are a bit higher this month, especially the number of files. At least some of the deposits must have been large, containing an unusually large number of files.
In addition to the usual suspects - plain text, stat package formats, MS Word, PDF - we have a very large number of unidentified files this month (2400+ application/octet-stream), and we also have a very small number of interesting formats (images, photoshop).
Monday, October 3, 2011
All Things Confidential
Tech@ICPSR will be at the University of Michigan's Michigan Union this Thursday to participate in the 2011 biennial meeting of our Organization Representatives from across the world of higher-education. We'll be batting lead-off for the All Things Confidential session at 9am EDT.
If you cannot attend the meeting in person, you can still attend virtually.
If you cannot attend the meeting in person, you can still attend virtually.
Labels:
announcement,
icpsr,
information technology,
restricted data,
technology
Location:
330 Packard St, Ann Arbor, MI 48104, USA
Saturday, March 26, 2011
February Deposits at ICPSR
February 2011 deposits (and their file formats) at ICPSR:
Nothing too exciting this month.
[ I thought I had posted this weeks ago, but clearly not. ]
| # of files | # of deposits | File format |
| 1 | 1 | application/msoffice |
| 284 | 23 | application/msword |
| 11 | 4 | application/octet-stream |
| 104 | 25 | application/pdf |
| 7 | 5 | application/vnd.ms-excel |
| 24 | 10 | application/x-sas |
| 80 | 28 | application/x-spss |
| 26 | 4 | application/x-stata |
| 2 | 1 | application/x-stuffit |
| 2 | 2 | application/x-zip |
| 19 | 6 | message/rfc8220117bit |
| 15 | 3 | text/html |
| 9 | 2 | text/plain; charset=iso-8859-1 |
| 2 | 2 | text/plain; charset=unknown |
| 122 | 21 | text/plain; charset=us-ascii |
| 9 | 5 | text/rtf |
| 1 | 1 | text/x-c; charset=iso-8859-1 |
| 22 | 7 | text/x-c; charset=us-ascii |
| 1 | 1 | text/xml |
Nothing too exciting this month.
[ I thought I had posted this weeks ago, but clearly not. ]
Monday, November 16, 2009
ICPSR Content and Availability

Legend:
Blue = Archival Storage
Yellow = Access Holdings
Green = both Archival Storage and Access Holdings
Red Outline = Web-delivered copy of Access Holdings
We're getting close to the one-year anniversary of the worst service outage in (recent?) ICPSR history. On Monday, December 28th, 2008 powerful winds howled through southeastern lower Michigan, knocking out power to many, many thousands of homes and businesses. One business that lost power was ICPSR.
No data was lost, and no equipment was damaged, but ICPSR's machine room went without power nearly until New Year's Day. In many ways we were lucky: The long outage happened during a time when most scholars and other data users are enjoying the holidays, and there was no physical damage to repair. The only "fix" was to power up the equipment once the building had power again.
However, this did serve as a catalyst for ICPSR to focus resources and money on its content delivery system, and therefore on its content replication story too. Some elements of the story below predate the 2008 winter storm, but many of the elements are relatively new.
ICPSR manages two collections of content: archival storage and access holdings.
Archival storage consists of any digital object that we intend to preserve. Examples include original deposits, normalized versions of those deposits, normalized versions of processed datasets, technical documentation in durable formats such as TIFF or plain text, metadata in DDI XML, and so on. If a particular study (collection of content) has been through ICPSR's pipeline process N different types, say due to updates or data resupplies, then there will be N different versions of the content in archival storage.
Access holdings consist of only the latest copy of an object, and often include formats that we do not preserve. For example, while we might preserve only a plain text version of a dataset, we might make the dataset available in contemporary formats such as SPSS, SAS, and Stata to make it easy for researchers to use. Anything in our access holdings would be available for download on our Web site, and therefore doesn't contain confidential or sensitive data. Much of the content, particularly more modern files, would have passed through a rigorous disclosure review process.
The primary location of ICPSR's archival storage is a EMC Celera NS501 Network Attached Storage device. In particular, a multi-TB filesystem created from our pool of SATA drives provides a home for all of our archival holdings.
ICPSR replicates its archival storage in three locations:
- San Diego Supercomputer Center (synchronized via the Storage Resource Broker)
- MATRIX - The Center for Humane Arts, Letters, & Science Online at Michigan State University (synchronized via rsync)
- A tape backup system at the University of Michigan (snapshots)
Some of our content stored at the San Diego Supercomputer Center - a snapshot in time from 2008 - is also replicated in the Chronopolis Digital Preservation Demonstration Project, and that gives us two additional copies of many objects.
An automated process compares the digital signature of each object in archival storage and compares it to a digital signature calculated "on the fly." If the signatures do not match, the object is flagged for further investigation.
The primary location for ICPSR's access holdings is also the EMC NAS. But in this case, the content is stored on a much smaller filesystem built from our pool of high-speed, FC disk drives.
ICPSR replicates its access holdings in five locations:
- San Diego Supercomputer Center (synchronized via the Storage Resource Broker)
- A tape backup system at the University of Michigan (snapshots)
- A file storage cloud hosted by the University of Michigan's Information Technology Services
- An Amazon Web Services (AWS) Elastic Computing Cloud (EC2) instance located in the EU region
- An Amazon Web Services (AWS) Elastic Computing Cloud (EC2) instance located in the US region
The AWS-hosted replica has been used twice so far in 2009. We performed a "lights out" test of the replica in mid-March, and we performed a "live" failover due to another power outage in May. In both cases the replica worked as expected, and the amount of downtime was reduced dramatically.
And, finally, our access holdings and our delivery platform are available on the ICPSR Web staging system. But because the purpose of this system is to stage and test new software and new Web content, this is very much an "emergency only" option for content delivery.
Labels:
icpsr,
infrastructure,
proposals,
technology
Monday, October 12, 2009
OR Meeting 2009 - Live Chat with Bryan Beecher and Nancy McGovern
Nancy McGovern and I co-hosted a "live chat" session at this year's meeting for Organizational Representatives (ORs). The video content of this is pretty light - just a few slides I put together to help generate discussion.
You can also find this session - and many more - on the ICPSR web site: http://www.icpsr.umich.edu/icpsrweb/ICPSR/or/ormeet/program/index.jsp.
Wednesday, September 2, 2009
ICPSR: Then and Now: Servers

In 2002 ICPSR had two main systems - a pair of Sun E3500s with 4GB of memory. One machine served as our production web server, and the second did everything else: general-purpose computing for data processing, Oracle database service, file service (NFS and CIFS via samba), DNS service, etc. We also had a very small number of additional machines, such as a system for testing new web applications. All of the machines were built by Sun Microsystems, used Sun's SPARC processors, and ran Sun's operating system, Solaris. We entered into a maintenance contract with Sun in case either of the machines had a problem, and my recollection is that it ran around $15k/year to cover the two big machines plus a handful of external storage arrays. To Sun's credit they were very solid machines.
In 2009 ICPSR has more servers than I can describe easily in a blog post. We still have a pair of machines for delivering web content and general-purpose computing, but they were built by Dell, use Intel processors, and run Red Hat Linux. Today's machines have much more memory and many more processors, and they too have been solid. But we also have many smaller machines with very specific roles: delivering network services (DNS, DHCP, etc); operating our LOCKSS network; staging new web content; replicating services for the Minnesota Population Center; hosting MySQL and Oracle databases; and so on. And, of course, in 2009 Sun Microsystems is about to be swallowed by Oracle.

However, this proliferation of server computing systems has likely reached its apogee at ICPSR. With the rise of virtualization and particularly the rise of the cloud, we're much more likely to build future systems in Amazon's Elastic Computing Cloud (EC2) rather than building them on real (or virtual) machines at ICPSR. For every rack-mount server we have at ICPSR, we probably have one much smaller blade server, and for every blade server, we probably have one EC2 instance running in the cloud.

My sense is that we'll continue this trend, and that where practicable, we'll deploy new systems in a cloud environment rather than purchasing new hardware. In addition to Amazon's cloud offering, the University of Michigan is deploying its own virtualization service, and that will be an attractive choice for systems that consume a lot of network I/O. Amazon charges for network I/O, but U-M does not.
We may also replace several virtual machines in the cloud with an out-sourced service: we already use SalesForce.com as our platform for managing "data leads." It's easy to imagine us adopting OpenID via a service provider such as RPX rather than hosting our own service locally or in a cloud, for example,
Thursday, July 23, 2009
New ICPSR web site
An "open beta" of our new web site is available. Here's a link to the site:
http://staging.icpsr.umich.edu/

In addition to the new look and feel, we've also made significant changes "under the hood." Perhaps the two biggest changes are with our search technology and with our overall technology platform.
Our new search technology is Solr, the search engine build on top of Lucene, which is a component of the Apache Project. Solr has all of the features and conveniences one expects in a modern search capability, and is a significant upgrade from the Autonomy Ultraseek product we have been using since the early 2000's.
Our new technology platform is largely Java servlets and JSP. While we still have many significant systems (e.g., our deposit submission system) on our legacy platform (perl CGI scripts), we'll build new systems on this new platform. We've been impressed with the quantity and quality of tools and supporting technologies that play well with Java, and by using JSP we also make it easier for the software developers and web designers to work together.
http://staging.icpsr.umich.edu/

In addition to the new look and feel, we've also made significant changes "under the hood." Perhaps the two biggest changes are with our search technology and with our overall technology platform.
Our new search technology is Solr, the search engine build on top of Lucene, which is a component of the Apache Project. Solr has all of the features and conveniences one expects in a modern search capability, and is a significant upgrade from the Autonomy Ultraseek product we have been using since the early 2000's.
Our new technology platform is largely Java servlets and JSP. While we still have many significant systems (e.g., our deposit submission system) on our legacy platform (perl CGI scripts), we'll build new systems on this new platform. We've been impressed with the quantity and quality of tools and supporting technologies that play well with Java, and by using JSP we also make it easier for the software developers and web designers to work together.
Monday, March 30, 2009
Blog grand opening!
Exploration. Evaluation. Deployment. Management. ICPSR is constantly looking at developing technologies, implementing the most promising ones for new projects, and surveying the landscape for new ways of managing our existing systems. It's a very busy place, and always very interesting.
I'm intending to use this space to let people know what's going on with technology at ICPSR, and hoping that people will find it interesting and informative.
I'll be spending a lot of time over the next month or two working with the Fedora, looking at ways of storing social science data and documentation in this open-source repository software. ICPSR already has a rich system of content delivery and existing preservation systems, and exploring the best way of integrating (or replacing) them with Fedora will be one of the first tasks to undertake.
I'm intending to use this space to let people know what's going on with technology at ICPSR, and hoping that people will find it interesting and informative.
I'll be spending a lot of time over the next month or two working with the Fedora, looking at ways of storing social science data and documentation in this open-source repository software. ICPSR already has a rich system of content delivery and existing preservation systems, and exploring the best way of integrating (or replacing) them with Fedora will be one of the first tasks to undertake.
Labels:
digital preservation,
fedora,
icpsr,
technology
Subscribe to:
Posts (Atom)
