Wednesday, October 19, 2011

The RCS becomes the DARS

ICPSR first launched its Restricted Contract System (RCS) more than two years ago.  Since that initial launch we've learned a lot:  who actually uses the system to apply for access to data; how they experience the system; how ICPSR contract administrators use the system; and, how to build in workflow to make it a smoother experience for all parties.

We relaunched the RCS last week, but with a new name:  the ICPSR Data Access Request System (DARS).  I suspect a lot of us will continue to call it the RCS, but it is the same system, but with a very different look and flow.

The DARS home page is the same as the old RCS system, and the most typical access method for initial use is from the home page of a study.  If a version of a study is available through a data-use agreement, then a link appears on its home page, and clicking that link navigates the visitor to the DARS.

Once there the visitor can initiate the data-use agreement process, going through the same general steps as before.  However, it is now much clearer when the agreement process has been completed, and the ball is now in ICPSR's court for review.  We've also worked hard to distinguish between essential elements of the agreement (e.g., if it changes, then the agreement must be reviewed and signed again), and which are more tangential (e.g., if it changes, ICPSR will be notified, but the agreement need not be signed off on again).

One element of the redesign is an explicit acknowledgement that this system may be used for any data-use request, and is not limited to only restricted-use requests.  We based this change on feedback from an internal team of reviewers who thought that the system should be able to work for any type of content that requires an agreement, even if it isn't particularly sensitive or confidential.

This design also recognizes that the applicant using the system may not necessarily be the PI who is requesting access to the data.  (In fact, we suspect that most applicants are not the PI.)  We therefore built views and rules to make it easier for, say, a population center data coordinator who may be working on several request for several PIs to get a better view of status across all requests.

Monday, October 17, 2011

Tech@ICPSR heads to Gartner

Tech@ICPSR is attending the Gartner Symposium/ITxpo this week in sunny Orlando.  Rather than writing a single blog post summarizing the event, I'll try to post a series of very short posts throughout the week as I attend different sessions.  (Something bigger than a tweet, but not as long as a typical blog post here.)

Friday, October 14, 2011

TRAC: A1.1: Mission statement

A1.1 Repository has a mission statement that reflects a commitment to the long-term retention of, management of, and access to digital information.

The mission statement of the repository must be clearly identified and accessible to depositors and other stakeholders and contain an explicit long-term commitment.

Evidence: Mission statement for the repository; mission statement for the organizational context in which the repository sits; legal or legislative mandate; regulatory requirements.



Here is our mission statement: 

ICPSR provides leadership and training in data access, curation, and methods of analysis for a diverse and expanding social science research community.

The important word in the context of TRAC requirement A1.1 is curation.  My recollection is that as we were crafting the mission statement, it originally contained many, many more words, and one of the most time-consuming tasks was for the group to pare it down.  I think the term curation was meant to imply not passive storage of content, but active management of our content, beginning with its receipt in the deposit system, and then onward.

Along with a mission statement we also have a strategic plan (which is getting a little long in the tooth now that it is 2011), and the commitment to long-term retention and management of digital information also appears in several places therein.

Wednesday, October 12, 2011

A very brief introduction to FLAME - ICPSR's File-Level Archival Management Engine

We've closed the books on 2011 Q3 and have moved on to the list of priorities for Q4.  One of the top priorities is a new project called FLAME (File-Level Archival Management Engine).

The goal of the project is to re-tool ICPSR's primary technology infrastructure so that is file-oriented rather than "study"-oriented.  This is essential to ICPSR's future for two main reasons.

One, more and more of our content doesn't fit nicely into ICPSR's classic "survey data and codebook" model.  We're starting to handle content like classroom observation sessions (video) and open-ended textual content (qualitative data), and even some of our existing content (TIGER/Line files, Census 2000 summary files, CCEERC reports) does not fit into the current object model without much contortion.

Two, the file-level is a much better fit for mapping business functions to the Open Archival Information System (OAIS) reference model (pink book), and for conforming to best practices, such as the Trustworthy Repositories Audit and Certification checklist.  For example, if we want to be able to demonstrate trustworthiness when it comes to the mapping from a file we deliver on the web site to a file we have in archival storage to a file that was deposited, we need to collect information and manage content at the file level.

I've been looking at the wiki for Archivematica, a site that I learned about from Nancy McGovern.  They've created a use-case and one or more related microservices for many of the boxes and connectors in the OAIS reference model.  I like the idea of linking the software directly to the OAIS reference model like this, and I'm intending to make great use of the Archivematica work to help us here.


Clip art credit: http://www.flickr.com/photos/bycp/5690269952/sizes/s/in/photostream/

Monday, October 10, 2011

September 2011 deposits

The report for September:

# of files# of depositsFile format
11application/msaccess
3714application/msword
132application/octet-stream
12528application/pdf
3715application/vnd.ms-excel
74application/x-sas
7727application/x-spss
144application/x-stata
11application/x-zip
41message/rfc8220117bit
44text/html
22text/plain; charset=iso-8859-1
44text/plain; charset=unknown
19630text/plain; charset=us-ascii
43text/rtf
11text/x-c; charset=unknown
54text/x-c; charset=us-ascii

Pretty typical formats, and pretty normal volumes.  A few deserve investigation (octet-stream), and a few plain text files have been tagged as C source code (as usual).

Friday, October 7, 2011

ICPSR wins grant from the Bill and Melinda Gates Foundation

ICPSR and partners at the University of Michigan received a grant from the Bill and Melinda Gates Foundation recently.  You can find the official link here.

 The link says that the grant is to house and make available to qualified researchers the data collected by the Measures of Effective Teaching project, and than means that you'll soon see yet another ICPSR operated web portal.

In addition to building and operating the portal, the project also requires us to process, preserve, and deliver quantitative data related to the Measures of Effective Teaching (MET) project, which is right in ICPSR's wheelhouse, of course.  The really new element for us, though, is the collection of video and "artifacts" related to the video.

ICPSR will use its existing Restricted Contract System (RCS) to screen applicants who want to access the video collection.  If approved the applicant will be able to access a video streaming server to view the videos in the collection.  The applicant will also be able to access the quantitative data in our Virtual Data Enclave (VDE).

The video collection is large compared to our current holdings of survey and government data.  My sense is that our collection will pretty much double in size, approaching 20TB total.  That's very big for us, but not nearly as big as some collections of video.


Wednesday, October 5, 2011

Web availability through September 2011

ICPSR web availability through 9/2011
Web site availability was good, but not great in September.  We've found that our Solr search query process is the most fragile piece of the infrastructure, and it got "stuck" on Sunday evening, 9/25.  Usually these are easy-to-correct faults; we just restart the tomcat instance hosting the Solr search query service.  But on this particular night the on-call missed the page, and the U-M Network Operations Center (NOC) did not open a ticket and phone the on-call, and so it lasted closer to 90 minutes.

During that time the web site was still usable, of course, and lots of functions would have worked normally (viewing pages, download studies, using SDA, etc).  But we start our "unavailability counter" whenever any part of the infrastructure is unavailable.  But my apologies if you were trying to search our catalog at that time.

Our analysis is that the virtual machine is running out of memory on our current (but old) web server.  We have a new 64-bit machine with significantly more memory available, and we'll been prepping it to take over for the old machine.  In the process of building the new machine we've been upgrading versions of Red Hat, tomcat, java, and many other key elements, and this has made the going a bit more slow than usual, but should give us a machine with better software.  And software that doesn't need to be upgraded right away (I hope!).