Friday, December 16, 2011

TRAC: A3.5: Seeking and acting on feedback

A3.5 Repository has policies and procedures to ensure that feedback from producers and users is sought and addressed over time. 

The repository should be able to demonstrate that it is meeting explicit requirements, that it systematically and routinely seeks feedback from stakeholders to monitor expectations and results, and that it is responsive to the evolution of requirements.

Evidence: A policy that requires a feedback mechanism; a procedure that addresses how the repository seeks, captures, and documents responses to feedback; documentation of workflow for feedback (i.e., how feedback is used and managed); quality assurance records. 



I think the market economy that keeps ICPSR in business is the very best evidence that the organization seeks input from its community, and applies that feedback to its operations, content selection, preservation strategies, and nearly every element of its business.

In practice we can see many different types of feedback mechanisms:  contract renewals; annual membership renewals; biennial Organizational Representative meetings; regular ICPSR Council meetings; and, regular participation at all sorts of public forms about social science research data, digital preservation, technology, etc.  It also happens electronically via social media, a helpdesk where a real person answers the phone and emails, and feedback pages on the main web portal.

In some ways it feels as if this TRAC requirement is aimed at organizations that might be funded by one community, like a national government, but used by a very different community, such as research.  In a scenario where the consumers and the payers are different, it is indeed critical that there be some mechanism to collect input, or the repository could enter a kind of "zombie" state where it ceases to serve its community effectively, but the funding organization continues to fund the repository nonetheless.

That said, I do think there is room for improvement in this area for ICPSR.  In particular, I think there is a great opportunity to work more closely with the individual data producers, engaging them in the curation process, and making the workflow - from deposit through eventual release - more transparent.

Wednesday, December 14, 2011

Google Music keeps the tunes playing

I started using the new Google Music production service.  I hadn't explored Google's previous offering, the Music Beta, all that much, but decided the time was right to dip a toe into the water.

The service has a lot of similarities to iTunes, of course, except one's library is in the cloud rather than on a PC (assuming one isn't using Apple's iCloud).  Google gives one free space to store 20k songs.  I'm using about 1% of that quota so far.

I like the idea of having a copy of our music in the cloud as an additional backup (or preservation copy), and it is also nice being able to use a standard browser window to manage and play the music.  One complaint I have about iTunes is that because it is conventional desktop software, one has to update it from time to time.  And this is somewhat more burdensome if one has to switch from a "standard" type of login on Windows to one with administrative rights, and then switch back again.

Google provides a tool which will copy music from one's existing storehouse (mine was an iTunes library).  The tool worked well for this purpose, and it did NOT require any administrative rights on my home WinXP (I know, I know) to download, install, and execute.  I started the copy one evening, and some 400 songs had been copied into Google Music by the morning.  One feature request:  It would be fabulous if the Music Manager tool would pull songs directly from a CD.

On the back-end I wonder if Google is using some form of de-duplication to minimize the amount of storage it needs to provision for this service?  It must be the case that there would be great overlap between music collections, particularly with the most popular songs, artists, albums, etc.  Google does such a good job of squeezing storage efficiency out of GMail; would expect them to do the same for their music service.

Monday, December 12, 2011

ICPSR web availability through November 2011

Now that's what I'm talking about!  (Click the image to see a more readable chart.)

ICPSR FY 2012 web availability through November 2011
After some shamefully low availability numbers in September and October, we've rebounded nicely in November (over 99.9%).

We saw two main problems in the month. 

One was a short outage where our search engine (Solr) faulted and required a restart.  We think this is due to a memory leak in Solr, and we are hoping that we can avoid the problem more completely once we move from our older 32-bit web hardware to our new 64-bit machine.  I'm hoping this happens before the end of the calendar year.

The other outage was due (we think) to a campus power blip that seemed to cause a fault with our EMC storage appliance.  While ICPSR never lost power, and while the machine room has a large, new UPS system, we speculate that the EMC got confused when it lost contact with UMROOT Windows Domain Controllers across campus due to the network path fluctuating.  The problem solved itself after 15 minutes, and it was the only anomaly that was coincident with the EMC hanging.

Friday, December 9, 2011

TRAC: A3.4: Formal, periodic review

A3.4 Repository is committed to formal, periodic review and assessment to ensure responsiveness to technological developments and evolving requirements.

Long-term preservation is a shared and complex responsibility. A trusted digital repository contributes to and benefits from the breadth and depth of community-based standards and practice. Regular review is a requisite for ongoing and healthy development of the repository. The organizational context of the repository should determine the frequency of, extent of, and process for self-assessment. The repository must also be able to provide a specific set of requirements it has defined, is maintaining, and is striving to meet. (See also A3.9.)

Evidence: A self-assessment schedule, timetables for review and certification; results of self-assessment; evidence of implementation of review outcomes. 



Steve Abrams from the California Digital Library gave an interesting talk earlier this year about the notion of applying a Neighborhood Watch metaphor to digital archives.  You can find a PDF of the slideshow here.

This is a nice paradigm, and it fits well with some of the work ICPSR is doing with its Data-PASS partners.  We're using the Stanford Lots of Copies Keep Stuff Safe (LOCKSS) software in a Private LOCKSS Network (PLN) to build a distributed archival storage network.  And in addition to the PLN, we have also built tools to verify the integrity of the PLN and its content.  We call this additional layer the SAFE-Archive, and the development has been led by the Odum Institute at the University of North Carolina.

I also see ICPSR periodically assess itself on a regular basis in response to opportunities to expand its reach thematically or technologically.  For example, as ICPSR enters into the world of digital preservation for video as part of two recent grants from the Bill and Melinda Gates Foundation, this drives ICPSR to re-evaluate how it manages content.

I'm not sure that these types of activities are as formal as the TRAC requirement might like, and so the action item might look more like a documentation project rather than adding a new activity into ICPSR's standard operating procedures.

Wednesday, December 7, 2011

November 2011 deposits at ICPSR

Chart?  Chart.[1]

# of files# of depositsFile format
11application/msaccess
3714application/msword
143application/octet-stream
16124application/pdf
161application/postscript
44010application/vnd.ms-excel
11application/vnd.ms-powerpoint
1061application/x-arcview
531application/x-dbase
102application/x-dosexec
11application/x-executable, dynamically linked (uses shared libs), not stripped
11application/x-rar
209application/x-sas
81application/x-sharedlib, not stripped
41application/x-shellscript
966733application/x-spss
234application/x-stata
53application/x-zip
121image/gif
11image/jpeg
21image/x-xpm0117bit
22message/rfc8220117bit
39712text/html
118text/plain; charset=iso-8859-1
138text/plain; charset=unknown
148242text/plain; charset=us-ascii
163text/rtf
41text/x-c++; charset=us-ascii
82text/x-c; charset=us-ascii
11text/xml

A relatively heavy month for deposits, this November 2011.  The number of SPSS files deposited is really impressive, and look to be the result of a small number of deposits, but with many data files.  Quite a bit of Excel too; more than SAS and Stata combined. 

We also have the usually fishy looking items that have been auto-detected as C or C++ source code, and which are actually text/plain (I suspect).  If the setups for a stat package contain the right types of comments in the right places, file is easy to fool.

The dBase and ArcView files are an interesting add to this month's listing.  We don't see too many of those.

[1]  This is a very small homage to mgoblog.

Monday, December 5, 2011

Collaborators, not depositors

ICPSR should stop accepting deposits.

Instead ICPSR should be recruiting collaborators.

To be sure ICPSR receives a great deal of its content via US Government agencies who have decided to outsource the digital preservation of their content to a trustworthy repository like ICPSR.  In this case the relevant contract, grant, or inter-agency agreement makes it clear what content will be coming to ICPSR to be curated and preserved.  In some cases the agency has little interest in depositing content ("Isn't that what we pay you for?"), and so the formal act of depositing content falls to the ICPSR staff anyway.

However, we also receive a considerable volume of content through our web portal where the depositor is external.  Sometimes we have worked hard to acquire the content, and the deposit is one milestone on a very long road, but other times the content comes to us unsolicited.  (I like to call these "drive-by deposits.")

In some cases the depositor is quite eager and able to help ICPSR with much of the curation work:  drafting rich descriptive metadata; organizing survey data and documentation into coherent groups; packaging other types of content into logical bundles (such as with our Publication-Related Archive); and, reviewing the data for possible disclosure risks.  Depositors may have access to resources like graduate students who can help with these tasks, and if the depositor is also the data producer, then s/he has valuable, unique insight into the data and documentation.  Unfortunately ICPSR is not well poised to tap into that expertise and those resources.

What would it take to get there?

ICPSR could separate the transactional step of submitting content (i.e., file upload concurrent with signature) from the iterative step of preparing metadata applicable to the submitted content.  In fact, one could even prepare metadata well before the submission transaction if the data producer had the interest and resources to prepare that information, but was not quite ready to share the data yet.  And, it would be equally permissible to submit the data for preservation and sharing, and then build the metadata slowly during the weeks and months following the upload.

If the data producer could also export the metadata in machine actionable formats, say, DDI XML for content which maps well to the classic "study" object that ICPSR has curated and preserved for decades, then there may be additional value to the producer. And introducing the structure that comes along with an XML schema like DDI might also be valuable to the producer in terms of thinking about and organizing the documentation, even for his/her own use.

In this world the ICPSR deposit system becomes a much shorter, much simpler web application.  And the ICPSR data management infrastructure would need to be opened up -- but with serious access controls -- so that content providers could access, create, and revise their documentation and metadata.  But the best thing about this world is that ICPSR gains a lot of collaborators, some who would be quite eager to work with us, I think.

Friday, December 2, 2011

TRAC: A3.3: Permission to preserve

A3.3 Repository maintains written policies that specify the nature of any legal permissions required to preserve digital content over time, and repository can demonstrate that these permissions have been acquired when needed. 

Because the right to change or alter digital information is often restricted by law to the creator, it is important that digital repositories address the need to be able to work with and potentially modify digital objects to keep them accessible over time. Repositories should have written policies and agreements with depositors that specify and/or transfer certain rights to the repository enabling appropriate and necessary preservation actions to take place on the digital objects within the repository.

Because legal negotiations can take time, potentially slowing or preventing the ingest of digital objects at risk, a digital repository may take in or accept digital objects even with only minimal preservation rights using an open-ended agreement and address more detailed rights later. A repository’s rights must at least limit the repository’s liability or legal exposure that threatens the repository itself. A repository does not have sufficient control of the information if the repository itself is legally at risk.

Evidence: Deposit agreements; records schedule; digital preservation policies; records legislation and policies; service agreements.



ICPSR has a standard agreement that is uses for all deposits.  This agreement grants ICPSR the non-exclusive right to replicate the content for preservation purposes and to deliver the content on our web site.  This language resides inside of our Deposit Form web application.

This works very well for deposits that come from a known source, such as a government agency with whom we have an agreement to preserve and deliver content, or an individual researcher with whom we have been corresponding.  In this case we have a good sense for who the depositor is, the role they play with regard to the data, and the mechanisms by which we can contact him/her.

Things become a bit messier what I will call a "drive-by deposit."  This is an unsolicited, unexpected deposit, and in this case the depositor agrees to give us permission to make copies of the content for digital preservation purposes and to deliver the content via our web portal.  That said, ICPSR does not require strong identities to execute a deposit, and so one could ask the question:  How does ICPSR know that the depositor himself/herself has the authority to grant us rights to preserve and redistribute the content?