Showing posts with label cyberinfrastructure. Show all posts
Showing posts with label cyberinfrastructure. Show all posts

Wednesday, July 17, 2013

ICPSR Web Availability - 2012-2013

Here are the final numbers for ICPSR's web site availability over our last fiscal year:

Click to embiggen
The year did not start off so well, and we reached the nadir quickly.  August 2012 was our worst period of availability in a very bad year for us overall.  January, March, and June 2012 also had very poor numbers.

The main antagonist we faced was a new and unusual problem with our Oracle database server.  For many years we would export the content for backup purposes each evening, and it worked well for a decade.  However, suddenly in 2012 we began to experience an outage just AFTER each export.  Despite intensive analysis by ourselves and local Oracle exports, we never could isolate the root cause of failure.

We eventually "solved" the problem by exporting our database only once per week v. once per day.  That left us more exposed to loss, of course, but it seemed to limit the outages to once per week v. once per day.

We then replaced the hardware with a new machine with a bit more processor and memory, but with blindingly fast solid-state drives. With the new machine deployed we returned to our daily export schedule, and the machine -- and our web availability -- have been in pretty good shape ever since. The machine went into service in April 2012, and the chart above makes it clear that life has been a little less hectic for our on-call engineer since then.

Monday, September 24, 2012

DuraSpace announces SDSC as storage partner

DuraSpace announced recently their relationship with SDSC as a storage provider for DuraCloud.  As I posted a while ago, we have been using both DuraCloud and their SDSC storage partner for a while.  It's great to see DuraCloud continue to grow.

Monday, September 17, 2012

Introducing ICPSR's Virtual Data Enclave (VDE)

The ICPSR Virtual Data Enclave (VDE) is a secure, virtual environment in which a researcher can analyze sensitive data, create research products, and then take possession of those products and analysis.  And while he VDE is not a substitute for a physical enclave and the types of security protocols it facilitates, the VDE is very much a potential substitute for the traditional practice of distributing confidential data via removable media, such as CD-ROMs.

The VDE uses much of the same technology that ICPSR uses internally for its Secure Data management Environment (SDE) which we have described a few times.  In brief, we use a virtual desktop environment that is operated by the University of Michigan's central IT shop and connect it to what we call our Private Network Attached Storage (NAS) appliance.  Both the virtual desktop and NAS are behind a firewall, and we use the firewall and Windows group policies to restrict what actions one pay perform.  Download?  Nope.  Cut-and-paste between the virtual desktop and the real desktop? Uh uh.  Capture screenshots by taking a picture of your monitor?  Well, ......

The virtual environment keeps sensitive datasets under lock and key at ICPSR, but makes it available to researchers.  The environment contains the usual array of applications used in the social sciences (but no email!), exactly the same sort of stuff we might set up for a visiting scholar or OR.

The researcher accesses the environment through a small, easy-to-download and -install client based on VMware View Client.  Authentication takes place using standard University of Michigan credentials which we (ICPSR) and others at UMich can issue to "friends."  Access between the real desktop and the virtual desktop is encrypted, and we are in the process of adding IPSEC encryption between the virtual desktop and the NAS.  (This latter traffic passes over UMich's data backbone, and access to those routers is limited to UMich central IT network engineers.)

The virtual machine is completely ephemeral and can be wiped after each use.  Any intermediate research or results are stored on the ICPSR NAS.  Our NAS is backed up weekly, and tapes are cycled off-site quarterly.  Once the research has been completed ICPSR retains a "just in case you need it" snapshot for up to three years.

Friday, July 20, 2012

Amazon's loss is SDSC's gain

One of the recent Amazon Web Services (AWS) power outages has left some of my EBS volumes in an inconsistent state.  If these were simple volumes, each containing a filesystem, then the fix is easy:  just dismount the filesystem, run fsck to check it, and then remount the filesystem after it has been fixed.  We have done this on several of our EC2 instances that had inconsistent volumes.

Unfortunately, for these particular volumes we have bonded them together to form a virtual RAID.  And this RAID is used as a single multi-TB filesystem which is much bigger than fsck can handle.  So we are kind of stuck.

One option would be to newfs the big filesystem, and to move the several TBs of content back into AWS, but that would be very slow.  And if there is another power outage......

So instead we called up our pals at Duracloud and asked them if they could help us enable replication of our content to a second provider.  (The first provider is - ironically - AWS.  But their S3 service, not their EC2/EBS service.)  They said they'd be happy to help, and, in fact, they will starting to replicate our content later this same week.  (Now that's service!)

The new copy of our content will now be replicated in...... SDSC's storage cloud.  This really brings us full circle at ICPSR since our very first off-site archival copy was stored at SDSC. Back then (like in 2008) it was stored in their Storage Resource Broker (SRB) system, and we used a set of command-line utilities to sync content between ICPSR and SDSC.  

The SRB stuff was kind of clunky for us, especially given our large number of files, our sometimes large files (>2GB), and our sometimes poorly named files (e.g., control characters in file names).  Our content then moved into Chronopolis from SRB, and then at the end of the demonstration project, we asked SDSC to dispose of the copy they had.  But now it is coming back......

Wednesday, June 6, 2012

ORCID: Open Researcher and Contributor ID

Nature ran a piece recently about ORCID.  ORCID stands for Open Researcher and Contributor ID, and has the goal of making it easier to know exactly who is responsible for a given bit of research.  ICPSR is a member of ORCID.

The article points out the most prolific researcher (by name) in 2011 is Y Wang who has published nearly four thousand times!  And, of course, this is because although Y Wang is a single name, it belongs to many different people.  And so how do you know how many of these publications we should credit to any one of those people known a Y Wang?

Having a unique way to identify someone is certainly useful, but it is mainly useful if EVERYONE participates in the scheme.  For example, what if all of the people who already have a ResearcherID (such as myself - E-9184-2010)  opt out of getting an ORCID too?  Or, will either of these be any better than using an email address to identify someone?

Wednesday, April 18, 2012

Great FLAMEing file identification service

Some parts of the FLAME project will lend themselves to a microservices approach.  Microservices, like cloud computing, is a trendy, useful concept, but without a crystal clear definition.  But my take is that a microservice is something that performs one small, but useful bit of work, and which can be swapped in and out of an overall architecture at a component level.  It needs to have very clear inputs and outputs, and cannot contain any "secret sauce" that isn't part of its functional role.

Do not try this street magic at home.
One common activity at ICPSR is automated file identification.  Historically we've done this with the venerable UNIX utility file, but where we modify the magic database heavily, particularly for the formats we see most often.  We also post-process the output from file where we need additional handling above and beyond the capabilities of the magic database (e.g., making decisions based on the name or extension of the file).

Managing the magic database is not for the faint of heart.  (Try updating the Vorbis section.)  And this management has gotten both harder -- RHEL 6 uses a new format for its magic database which is incompatible with RHEL 5 -- and easier -- the new format eliminates the pesky magic.mime database.  However, we've gotten reasonably competent at managing magic and have come to rely on it for file format identification.

In support of the FLAME project we even created a little web service that takes a file's content and its name as input, and delivers a little snippet of XML as the output.  The XML contains the "human readable" answer from our magic database and the "MIME type" too.  This is our first FLAME-inspired web service.

If you'd like to try it, you can use your favorite form-capable URL transfer utility to do so.  Here's an example where I have run curl on one of our RHEL machines:


dhcp-bryan:; curl -F "file=@uuid-comparison.xlsx;filename=uuid-comparison.xlsx" www.icpsr.umich.edu/cgi-bin/wsifile<?xml version="1.0" encoding="utf-8"?><wsifile><ifile>Microsoft Excel</ifile><ifilemime>application/zip; charset=binary</ifilemime><uploadInfo>application/octet-stream</uploadInfo></wsifile>

feeding in an Excel file as the input, and another with a plain text file:

dhcp-bryan:; curl -F "file=@/etc/resolv.conf;filename=resolv.conf" www.icpsr.umich.edu/cgi-bin/wsifile<?xml version="1.0" encoding="utf-8"?><wsifile><ifile>ASCII text</ifile><ifilemime>text/plain; charset=us-ascii</ifilemime><uploadInfo>application/octet-stream</uploadInfo></wsifile>

and an interesting MS Word file:

dhcp-bryan:; curl -F "file=@2011-03CouncilPandAminutes.doc;filename=2011-03CouncilPandAminutes.doc" www.icpsr.umich.edu/cgi-bin/wsifile<?xml version="1.0" encoding="utf-8"?><wsifile><ifile>CDF V2 Document, Little Endian, Os: Windows, Version 5.1, Code page: 1200, Number of Characters: 0, Name of Creating Application: Aspose.Words for Java 4.0.3.0, Number of Pages: 1, Revision Number: 1, Security: 0, Template: Normal.dot, Number of Words: 0</ifile><ifilemime>application/msword; charset=binary</ifilemime><uploadInfo>application/octet-stream</uploadInfo></wsifile>

Feel free to try it out, and to post reactions, suggestions here.

Wednesday, January 18, 2012

Disaster Recovery v. High Availability

A question I often receive from customers and colleagues is:  If ICPSR has a replica of its production delivery system in Amazon's cloud, why is it that the web site is sometimes down due to scheduled maintenance or unplanned outages?

The short answer is:  ICPSR's cloud replica serves a disaster recovery (DR) purpose, but not a high availability (HA) purpose.  Of course, more often than not, this generates a look that falls somewhere between Bah! and This sounds like some made-up IT nonsense!  However, it really is the answer.  But that begs the question:  What's the difference between DR and HA?  But first a trip back in time....

As some long-time ICPSR clients may recall, the ICPSR delivery system was off-line for nearly a week during the holiday break between 2008 and 2009.  The root cause was a long power outage due to a major ice storm in the Midwest which knocked out power to many homes and businesses, including many in Ann Arbor.  And because ICPSR resides in a building just a little bit off the University of Michigan's central campus, we're just like any other home or business that waits for DTE Energy to restore power.

As one might expect both myself and the ICPSR Director at the time, Myron Gutmann, were quite anxious for the power to be restored.  The storm had caused so much damage that it wasn't at all clear when the building's power would be restored.  And, after the first few days without power - and heat - the building's pipes were in danger of bursting.  Things were looking pretty bad.

However, as it turned out we had been experimenting with Amazon's new computing and storage cloud just prior to the storm.  It would be pretty easy to stand up a minimal web server in Amazon's cloud, something that would basically say Yes, we know our delivery system is down, and we're sorry about that.  And here's the best guess from the local power company about when power will be restored.  We then worked with some of our colleagues at the University of Michigan and the San Diego Computing Center to update the system that maps names (like www.icpsr.umich.edu) to network addresses so that ICPSR's URLs for its web site would point to this new, minimal web server in Amazon's cloud.  That didn't fix the problem, of course, but it let people know that ICPSR knew there was a problem, and shared the best information we had about the problem.

Once power was restored and the main delivery system came back on-line, I had a long conversation with Myron about how we wanted to position ICPSR for any future problem like this.  What if the building lost power again for an extended period?  What if a tornado knocked down the whole building?  What if the U-M suffered some catastrophic problem with its network?

One option was to change the architecture of ICPSR's delivery systems.  Rather than having a complex series of simple web applications, we could redesign and rebuild the whole system so that it would also contain a middle layer of technology that would catch and route incoming requests to one of many delivery system components.  And rather than having a single production system at the University of Michigan, we would build a multi-site production system spread across multiple network providers and service providers so that no single problem would disrupt services.  This is essentially the high availability (HA) version of ICPSR's delivery system.  It would have the virtue of providing true 99.99%+ reliability, but would cost plenty of money to design, build, and operate.  If you are running IT systems for a bank or a hospital or an aircraft carrier, you build them with HA.  But what about a data archive?

Another option was to keep the ICPSR delivery architecture the same, but replicate it somewhere off-site.  Automated jobs could keep the web content, data content, and web applications synchronized.  And an easy - but manual - process could be used to redirect traffic to the replica when needed.  In this world there would still be plenty of times where a component of ICPSR's delivery system might be off-line due to maintenance or a fault, but if the maintenance or fault was long-lived, then the replica could be pressed into service.  This type of solution would be inexpensive to design, deploy, and operate, and would deliver a credible disaster recovery (DR) story, but would probably only give us uptime somewhere between 99.0% and 99.9%.  Would that be good enough?

In the end, of course, we decided that the best use of resources would be to build a system that would still have some outages from time to time, but which would never again be off-line for an entire week.  We set an availability goal of 99.5% for each month across all components.  That is, every time a single component faults - search, download, online analysis, and so on - it counts against the uptime of the WHOLE system.  And we would leave it up to the judgement of the on-call engineer to decide when a problem was likely to be long-lived enough to warrant a switch to the replica.

So we chose DR instead of HA.

Looking back, my sense is that we made the right decision.  In practice we seem to hit our 99.5% availability goal most months, and because we did not tie up our software and systems development resources on rebuilding the delivery system to guarantee HA, we were able to design and build systems like our Restricted Contract System, Secure Data Environment, and Virtual Data Enclave.  Of course, when we need to perform a major bit of maintenance like last weekend where it is important that we continue to point www.icpsr.umich.edu at the production system rather than the replica, it always makes me wonder about the HA alternative.

Wednesday, January 4, 2012

Systems Architect Senior job posting @ ICPSR

We've posted another job opening on my team at ICPSR.  I'll include the text from the job description below, but the short-lived link to the U-M job site is http://umjobs.org/job_detail/65035/systems_architect_senior.

In brief we're looking for someone with deep experience building production operational environments for web applications who can apply those skills to the ICPSR environment.  We currently have a mix of stuff running on real hardware at ICPSR and virtual hardware in Amazon's EC2, and we have a mix of legacy Perl CGI code and newer Java-based web applications.  A mix of myself and Steve Burling (who retires at the end of January) have been the main architects of the environment, but it has grown so much over the past 5-10 years that it has become a full-time job.

Here are the details:


Systems Architect Senior

Job Summary

The Inter-university Consortium for Political and Social Research (ICPSR), the world's largest archive of digital social science data, is now accepting applications for a Systems Architect Senior. ICPSR is a unit within the Institute for Social Research at the University of Michigan. ICPSR's data are the foundation for thousands of research articles, reports, and books. Findings from these data are put to use by scholars, policy analysts, policy makers, the media, and the public. This position reports to the Assistant Director, Technology Operations, Computer and Network Services.

Responsibilities*

Responsibilities
This position is responsible for the design, implementation, maintenance, and regular management of ICPSRs web systems development, staging, production, and disaster recovery operational environments. This consists of several distinct platforms, including local physical hardware and virtual systems hosted in Amazons Elastic Computing Cloud (EC2). The successful candidate will also work closely with the Assistant Director, Software Development, Computer and Network Services to define and implement functional requirements.

One, this position will select, install, and manage integrated development environment (IDE) software on developer workstations, and the underlying software repository. The incumbent systems are Eclipse and CVS, respectively.

Two, this position will select, install, manage, and maintain the testing, staging, and production platform environments used by ICPSR to deploy and test new web applications. The incumbent web application server is Apache Tomcat, sometimes run as a stand-alone web server and sometimes as a client to Apache Httpd. The incumbent server platform is a mix of local, physical servers and Elastic Computing Cloud (EC2) instances running within Amazon Web Services (AWS). ICPSR has interest in exploring a more complete. cloud-based web application platform such as AWS Elastic BeanStalk.

Three, this position will manage the AWS-hosted replica of ICPSRs production web environment. This includes building and maintaining tools that synchronize software, static content, and database content between the production environment and the replica environment.

Four, the over-arching responsibility of this position is to maintain and improve ICPSRs capacity for delivering high-availability, high-performance web-based services, managing the tension between the desire to have well-defined, documented, predictable deployments and business processes and the desire to have fast moving, fluid, and flexible deployments and business processes.

Required Qualifications*

BS in Computer Science, Computer Engineering, or at least eight years of experience with designing and managing complex web application hosting environments
Two or more years of experience with J2EE application servers (such as Tomcat)
Two or more years of experience with virtualization products or services, such as Amazon Web Services
In-depth expertise with RedHat Enterprise Linux 5 and 6
In-depth knowledge of networking principles and network support
In-depth knowledge of Web technologies (Apache, Tomcat)
Experience operating an RDBMS (Oracle, MySQL)
Experience working with monitoring tools, control software (CVS , Subversion, Perforce), and build tools (Make , Ant)
Strong understanding of web application architectures
Enthusiastic self-starter who works well with other team members
Excellent inter-personal skills with the ability to communicate clearly to peers, vendors, customers, and colleagues

Desired Qualifications*

MS in Computer Science or Computer Engineering
At least five years of experience with J2EE application servers (such as Tomcat)
Experience with Veritas products (Veritas Netbackup)
Storage experience (EMC)
Experience as an Infrastructure Engineer in a high availability environment
Expertise working with monitoring tools , version control software (CVS , Subversion, Perforce), and build tools (Make , Ant)

Underfill Statement

This position may be underfilled at a lower classification depending on the qualifications of the selected candidate.

U-M EEO/AA Statement

The University of Michigan is an equal opportunity/affirmative action employer.

Wednesday, December 28, 2011

Starting the FLAME

In an earlier post I described a major new project at ICPSR called FLAME.  FLAME is the File-Level Archival Management Engine, and will become the new repository technology platform ICPSR uses to curate and preserve content.  As the name implies, the main molecule of information upon which FLAME will operate is a "file" which is different than the main molecules used at ICPSR today, the "deposit" and the "study."  In the big picture the activities at ICPSR will not change much:  we will still collect social science research data, curate them, preserve them, and make them available in a wide variety of formats and modes.  But when one looks at the details, an awful lot will change.

So when one is going to change everything, where does one start?

Fortunately we have a ready-made starting point with the Open Archival Information System (OAIS) reference model.  While this does not give us a blue print of what to build, it does give us a model to use as we construct our blueprints.  I believe this is very much what the folks at Archivematica have done.

So the question becomes:  How do we translate a high-level reference model that contains functions such as Receive Submission to the low-level blue prints one needs to reconfigure process and build software?  What kind of web applications do I need for Receive Submission?  What should they do?  Should that box that contains the submitter's identity be an email address?  A text string?  An ORCID?

So how to start?

One of my colleagues, Nancy McGovern, suggested we brainstorm 6-12 medium-level statements for each of the functions in the OAIS reference model.  We started with Receive Submission, and indeed generated 12 statements.  (The analogue at Archivematica is Receipt of SIP.)  One example is:

The producer provided basic provenance information at deposit

If the metaphor for building FLAME is building a house, then OAIS plays the role of high-level best practices.  The statements (like above) play the role of floor plans and elevations; those things to which most people can relate and make decisions.  So this is moving in the right direction, but we're still lacking the blueprints.

The next step is to take a statement like the one above and turn it into requirements for software (and for process).  One example requirement that flows from the statement above is:


FLAME should capture the following provenance information from the files after each content transfer:

i. Date and time at which each file is received
ii. Checksum of each file
iii. MIME type of each file
iv. Original name of each file
v. Packaging information (e.g., file was part of a Zip archive)

We can then discuss these low-level requirements with stakeholders, such as the acquisitions team, and with the technology team, such as a software developer who may have additional questions (e.g., "Well, what sort of checksum do you want - MD5, SHA-1, or something else?").

Right now we are working through the details of Receive Submission, and the next few stops on the roadmap will likely be in Ingest as well.  We're documenting both the high-level statements and the low-level requirements in a Drupal CMS that we use as our Intranet.

Wednesday, November 16, 2011

DuraCloud Archiving and Preservation Webinar

Shameless self-promotion alert...

The nice folks at DuraSpace have published the audio and video from the recent webinar that Michele Kimpton (CEO, DuraSpace) and I gave on DuraCloud.


Michele spends the first 5-10 minutes talking about the business case behind DuraCloud, and then I spend about 30 minutes talking about ICPSR and how we came to use DuraCloud to store a copy of our archival holdings.

Wednesday, November 9, 2011

ICPSR's Secure Data Environment overview

Jenna Tyson is a graphic artist on staff at ICPSR.  Over the fast few years Jenna has helped me out with displays for poster sessions, transforming the mediocre layout I produce with a true work of art.  I've posted some of her work here in the past.

I asked Jenna if she could create a logo for our Secure Data Environment (SDE), and above you can see the one that I liked best.  I leave it as an exercise to the reader to decide if the terrified individual in the picture is a defeated intruder or a frustrated ICPSR data curator.

The blog contains several posts that go into some detail about the software and security components behind the SDE, but I'm not sure that I ever posted a high-level description to set context, scope, and purpose.  And so along with Jenna's logo, I present the story behind the SDE.




The ICPSR Secure Data Environment (SDE) is a protected work area that uses technology and process to protect sensitive social science research data from accidental or deliberate disclosure.  The SDE exploits commonly used security technologies such as firewalls, ActiveDirectory group policies, and network segmentation to minimize unwanted access between the SDE and outside world.  Further, it takes advantage of work processes which require strict control of when data may be moved between the SDE and external locations.

Data enter the SDE through ICPSR's deposit system.  Depositors upload their content to a web application on our public web portal where it is encrypted.  An automated process "sweeps" content from the portal several times per hour, moving it to the SDE, where it is then unencrypted.  The content resides on a special-purpose EMC Network Attached Storage (NAS) appliance which services ICPSR's SDE.  The appliance uses private IP address space which is only routed within the University of Michigan enterprise network, and is also protected by a firewall.  Further, NAS shares are exported only to specific machines and only to specific ActiveDirectory groups.

ICPSR data managers must be located on the University of Michigan enterprise network to access the SDE.  (They may use the University of Michigan VPN client to access the network from remote locations, and this requires strong authentication and implements strong encryption.)  Data managers run a simple utility to "log in" to the SDE.  Once logged into the SDE they are assigned to a disposable virtual Windows 7 desktop system which is configured to persist any content on the ICPSR SDE NAS.  Any content stored on the virtual desktop system is destroyed once the image is terminated.

Data curators are not allowed to access the Internet or email within the SDE, and they do not have access to local system ports (e.g., USB).  Clipboards are NOT shared between the SDE and the local machine, and so there is no ability to "cut and paste" between the two environments.  It is possible, of course, for data curators to take notes from what they see on the screen, and to take screen snapshots, but ICPSR management considers these to be acceptable risks.

Data curators may release data from the SDE via two mechanisms.

One, they may submit completed work for release via an internal work system called turnover.  This queues material for placement in archival storage, and also queues related material for release on the web site.  A release manager reviews all content before allowing it on the web site.

Two, they may submit unfinished work for transfer outside of the SDE.  In this case a request appears in the inbox of the data curator's supervisor who may then review the request, and then accept or reject it.  If accepted the content is available to the data curator through a simple file retrieval mechanism, and the transfer is logged.

ICPSR has contracted the services of a "white hat" ethical hacker to assess the security vulnerabilities on the SDE.  ICPSR has already implemented small changes within the SDE based on preliminary reports from the contractor.

Wednesday, October 26, 2011

Using DuraCloud for Archiving and Preservation

I'll be joining Michele Kimpton, CEO of DuraSpace, on a webinar next Wednesday (November 2, 2011).  Our topic is DuraCloud, and how one can use this cloud-based service as part of one's digital preservation strategy.

I think sometimes people will view the cloud as an alternative to keeping and maintaining local copies, but at ICPSR we're using the cloud as an easy-to-manage storage location to supplement more conventional locations, such as local NAS storage and the University of Michigan's "Value Storage" service.

Here is a copy of the invite that went out via email:


DuraSpace

You are invited to attend the following event:

Using DuraCloud for Archiving & Preservation
Wednesday, November 2, 2011
1:00p.m. - 2:00p.m. Eastern Standard Time

Presented By:
Michele Kimpton, DuraSpace Chief Executive Officer & DuraCloud Project Director &
Bryan Beecher, Director of Computing & Network Services,Interuniversity Consortium for Political and Social Research (ICPSR)


Having a hard time keeping up with current preservation and archiving practices?
Are you finding the task of archiving your content complicated, costly and confusing?
Then you need to join us for a free webinar that details how DuraCloud can be part of your preservation and archiving solution.

This webinar will discuss how to use DuraCloud as a component of your archiving and preservation strategy. An overview of the service will include what it is, how it works, and the benefits it has to offer. Additionally, Bryan Beecher, Director of Computing & Network Services at ICPSR, will present ICPSR's preservation and archivingstrategy. Bryan will share how DuraCloud and other methods have been implemented to meet ICPSR's preservation and archiving goals.

If you are interested in attending please thoroughly complete the registration process (below) to receive your unique login url. Be sure to SAVE the return email you receive from Infinite Conferencing as it will include your unique login information.
 

Below is the call-in information for the event.
Via Skype (Free, World): Dial +9900827047086940 
Via Phone (Toll, US): Dial +1(201)793-9022 Enter Room Number: 7086940


The maximum capacity for this web seminar is 99 participants. The event will be recorded and slides will be available for viewing after the event at http://duraspace.org/web_seminars
.

Please contact Kristi Searle at ksearle@duraspace.org
 with any questions.

**Please be aware of Infinite Conferencing System Requirements:
-          Internet connection speed of 128 kbps or higher is recommended
-          Microsoft Windows XP, Vista, Windows 7, or Server 2003
-          Internet Explorer 6.0 SP2, 7.0, 8.0 & 9.0, Firefox 3.0x/3.5, 4, 5 and Chrome 12 browsers
-          Apple Mac with Intel CPU, Mac OS X 10.5/10.6, Safari 4.x, 5.x or Firefox 3.x, Java 1.5+
-          Linux, Unix, or Solaris with Mozilla 1.0+
-          Cookies and Scripting enabled in browser

Wednesday, October 19, 2011

The RCS becomes the DARS

ICPSR first launched its Restricted Contract System (RCS) more than two years ago.  Since that initial launch we've learned a lot:  who actually uses the system to apply for access to data; how they experience the system; how ICPSR contract administrators use the system; and, how to build in workflow to make it a smoother experience for all parties.

We relaunched the RCS last week, but with a new name:  the ICPSR Data Access Request System (DARS).  I suspect a lot of us will continue to call it the RCS, but it is the same system, but with a very different look and flow.

The DARS home page is the same as the old RCS system, and the most typical access method for initial use is from the home page of a study.  If a version of a study is available through a data-use agreement, then a link appears on its home page, and clicking that link navigates the visitor to the DARS.

Once there the visitor can initiate the data-use agreement process, going through the same general steps as before.  However, it is now much clearer when the agreement process has been completed, and the ball is now in ICPSR's court for review.  We've also worked hard to distinguish between essential elements of the agreement (e.g., if it changes, then the agreement must be reviewed and signed again), and which are more tangential (e.g., if it changes, ICPSR will be notified, but the agreement need not be signed off on again).

One element of the redesign is an explicit acknowledgement that this system may be used for any data-use request, and is not limited to only restricted-use requests.  We based this change on feedback from an internal team of reviewers who thought that the system should be able to work for any type of content that requires an agreement, even if it isn't particularly sensitive or confidential.

This design also recognizes that the applicant using the system may not necessarily be the PI who is requesting access to the data.  (In fact, we suspect that most applicants are not the PI.)  We therefore built views and rules to make it easier for, say, a population center data coordinator who may be working on several request for several PIs to get a better view of status across all requests.

Wednesday, October 12, 2011

A very brief introduction to FLAME - ICPSR's File-Level Archival Management Engine

We've closed the books on 2011 Q3 and have moved on to the list of priorities for Q4.  One of the top priorities is a new project called FLAME (File-Level Archival Management Engine).

The goal of the project is to re-tool ICPSR's primary technology infrastructure so that is file-oriented rather than "study"-oriented.  This is essential to ICPSR's future for two main reasons.

One, more and more of our content doesn't fit nicely into ICPSR's classic "survey data and codebook" model.  We're starting to handle content like classroom observation sessions (video) and open-ended textual content (qualitative data), and even some of our existing content (TIGER/Line files, Census 2000 summary files, CCEERC reports) does not fit into the current object model without much contortion.

Two, the file-level is a much better fit for mapping business functions to the Open Archival Information System (OAIS) reference model (pink book), and for conforming to best practices, such as the Trustworthy Repositories Audit and Certification checklist.  For example, if we want to be able to demonstrate trustworthiness when it comes to the mapping from a file we deliver on the web site to a file we have in archival storage to a file that was deposited, we need to collect information and manage content at the file level.

I've been looking at the wiki for Archivematica, a site that I learned about from Nancy McGovern.  They've created a use-case and one or more related microservices for many of the boxes and connectors in the OAIS reference model.  I like the idea of linking the software directly to the OAIS reference model like this, and I'm intending to make great use of the Archivematica work to help us here.


Clip art credit: http://www.flickr.com/photos/bycp/5690269952/sizes/s/in/photostream/

Monday, September 19, 2011

Convergence

The always entertaining and highly informative Moxie Marlinspike gave a very interesting talk at BlackHat USA 2001 about SSL.  This is the technology that (in theory) secures our communication channels on the Internet, keeping information like credit card numbers out of the hands of the bad guys.  I've seen past talks by Moxie where he describes the many flaws with SSL, but in this talk he introduces a new solution called Convergence.


The talk is fascinating, and I highly recommend watching it.  (It's on YouTube.)  It's about 45 minutes long, so enjoy over lunch.

In brief, Moxie cites two problems with the current SSL model, which requires all of us to trust Certificate Authorities (CAs), which have been hacked with increasing frequency, and which have also demonstrated drunk and disorderly behavior at times.  One, we have to trust them forever.  Two, there is no reasonably way to change who you trust.  For example, if one decided that Comodo (one of the largest CAs) just could not be trusted any longer, one could deleted Comodo from his/her browser's "trust database."  But doing this would make a large number of Internet web sites (20%) unusable.

Convergence replaces CAs with one or more self-selected "notaries" each of which can use a different method to ascertain whether a certificate is valid, including a self-signed certificate.  One may also use a "bounce notary" to separate those that know who you are from those that know where you are browsing.

Wednesday, August 24, 2011

E Amazon Unum - Out of Amazon, One

Photo from http://www.flickr.com/photos/polselli/1250189137/
Amazon announced a new "region" in the United States last week:  the GovCloud region.  The geographic location of the new region is on the US West Coast, but the logical location is Washington, DC.

Amazon says that ALL of its availability regions offer FISMA Moderate security controls, but this region offers one additional feature and demands one additional requirement so that it "supports the processing and storage of International Traffic in Arms (ITAR) controlled data and the hosting of ITAR controlled applications."  The post goes on to say that:
As you may know, ITAR stipulates that all controlled data must be stored in an environment where logical and physical access is limited to US Persons (US citizens and permanent residents). This Region (and all of the AWS Regions) also provides FISMA Moderate controls. This means that we have completed the implementation of a series of controls and have also passed an independent security test and evaluation. Needless to say, it also supports existing security controls and certifications such as PCI DSS Level 1, ISO 27001, and SAS 70.
This gets interesting for organizations like ICPSR that conduct a lot of business with the US Government.  Earlier this year we spent mounted a significant effort to categorize the security level for content stored in our archive, and then documented the relevant NIST security controls.  It is easy to imagine this this type of effort will repeat itself as we interact with more federal agencies, and as those agencies struggle to become compliant with FISMA.

However, if I can short-circuit the process by using Amazon Web Services as my "machine room," and relying on Amazon's existing certifications and controls, then I may be able to ease the burden of writing and maintaining (and possibly implementing!) our own controls.  I would not expect to eliminate the entire effort of documenting NIST security controls, but I may be able to point to Amazon's existing controls and documentation for, say, those controls related to the physical machine room.  And remote access.

Indeed, instead of an AWS-hosted instance creating a barrier to a project ("oh no, if we build this in the cloud, we'll need to re-do all of the relevant NIST controls!"), it would facilitate the project. 

Monday, July 25, 2011

An open letter to our storage system

Dear Storage System:

I think it is time that we had a talk.  Not a friendly chat, but a serious, heart-to-umm-bus talk.  About real issues.

You know, we used to have one of your older cousins live with us.  Your cousin's official name was just a bunch of letters and numbers, but we always said "the Naz" for short.  Kind of like "the Fonz."

The Naz was sweet.  Everyone loved the Naz.  When the Naz felt the least bit sick, in a flash there would be a call home, and then before you could say, "broken disk drive," some dispatcher would be on the phone with us, offering to bring out a replacement part that same day.  Sometimes we'd even get the phone call before we knew there was a problem.  Those were the times!

Sure, some people on campus made fun of us, but we loved the Naz anyway.  "Hey, why do you sepnd so much on storage for your Naz?  Don't you know you can buy storage from us way, way cheaper?"  Yes, they would say such things.  Other storage providers can be so cruel.

But we didn't care.  We loved our Naz.  Dependable, reliable, but, sure, a little expensive.  But we thought that we were spending our money wisely on the Naz.  Not a bit of trouble.  We loved our Naz.

But, as all good things go, so did the Naz.  Parts got old.  The cabinet got dusty.  Sure, a 300GB disk drive seemed big enough back in 2005, but now it just seemed, I don't know, quaint.  And, you know, the Naz never had a second data mover....

And, so, after a lot of planning, and a lot of work, we put the Naz out to pasture.  The Naz isn't storing and serving data like in the old days; it just mills about the machine room, chewing on electricity, and soaking up the air conditioning.  Visitors come by to visit, and they never stay with the Naz for very long.

Of course, as we were reading sweet bedtime stories to the Naz, we were grooming you for the job.  You were new and shiny.  We didn't know you very well, but we new our Naz, and we knew you were supposed to be just like our Naz.  But newer.  And better.  And with more blue lights.  And we were excited.

Things started well.  In many ways you were just like the Naz, but better.  Your disk drives were hefty.  You had newer software.  You even seemed a little faster, just like your car does after you wash it.  This was great.

But then the behavior problems started.  You know what I mean.

Like the time that you knew full well that there was a problem, but did you call home for help?  No.  You made us do it.  Why?  Why wouldn't you use the nice telephone line we left for you?

And, sure, after we finally convinced you to call home, you then told the foulest lies.

"We've tried to log in to your system, but the password doesn't work!" the techs would say.  Why did you give them the wrong password?  Do you think this is a game?

"Please download DiskDebunker v4.3 from our mirror site.  Install if on a Windows ME machine, configured for use on a private network, and use it to assess if the storage processor valve flanges are flush.  This will produce a 700GB file called DataGrommit.zip, and you should then upload that to the Easy Web support site." they would say other times.  I can only imagine the wild tales you must have told to confuse them so.  Why couldn't you just be honest and authentic, and tell them that one of your disk drives faulted, and all we needed was a replacement?  Why couldn't you be more like the Naz?

"I have read and understand all site messages." said the service requests.  Yeah, sure.  Why did our little troublemaker say this time?

This behavior must stop!

We're scheduling a little something that some people call "an intervention."  We want you to hear the problems that your mischief is causing.  We want you to hear the stories.  We want you to hear it from the Naz.

We don't want it to end this way.  We want it to work.  But we've had enough.  Don't make us call the people at Property Disposition to come get you.  You won't like where they will take you.

Let's give it one more try.

Your friends @ ICPSR.

Wednesday, July 13, 2011

ICPSR's Secure Data Environment (SDE)

ICPSR has designed, built, and deployed what we call the Secure Data-processing Environment (SDE) over the past twelve months.  This is a tightly managed, highly controlled environment in which many members of the ICPSR staff perform their day-to-day data management (data processing) work.

The main business requirement behind the SDE is that it should be difficult, if not impossible, for content to leak out without a member of the staff taking an explicit action, such as running a program which formally releases content on the web site and commits it to archival storage.  For example, it should not be possible for someone to upload a data file into a web form, or to attach it to a piece of email.

The design called for many changes to ICPSR's technology infrastructure.  We separated our storage into two pools - Private (accessed within the SDE) and Semi-Private (which is more accessible).  We separated our network into three main virtual LANs - Private, Semi-Private, and Public.  We also updated many, many software systems so that they would operate properly within the SDE.  And we also changed processes to conform to the new business requirements.  For example, if one process required a data processor to send an email containing a data file to someone else at ICPSR, we changed the process so that email was not required.

I'll post a series of articles over the next few weeks with more details about the SDE and its technology.  This will include posts about how we separated storage; how we segmented the network; how we used virtualization technology to solve certain problems; how we changed key software systems; and, how the SDE changed business processes at ICPSR, and how it continues to do so even today.

Wednesday, July 6, 2011

ICPSR web portal availability in 2010-2011

It's that time again:  the end of another fiscal year.  And that means it is also time for my annual summary of ICPSR web portal availability.

The leftmost month above is July 2010 and the leftmost is Jun 2011.  The vertical axis shows availability for each month in terms of a percentage.  Our goal is to hit or exceed 99% availability each month.
All in all it was a pretty good year for ICPSR's production web portal.  Our web portal hosts many different sites (ICPSR proper, NACJD, NACDA, SAMHDA, DSDR, CCEERC, the ICPSR Summer Program, and many more sites).  We were able to exceed 99.75% availability most months, and only had two months (January and June 2011) where our level was a bit lower.

The main culprit of downtime throughout fiscal year 2011 was due to defects in software.  As we have been retooling our technology environment from Perl and CGI scripts to Java applications, we have been making greater use of systems like Hibernate and Lucene.  My sense is that we're relying more and more on open source middleware, and while that has the advantages of making it easier to develop software quickly, it also means that a problem in the underlying middleware can affect our overall availability.  Some of this is due to buggy software; some is due to our learning curve on how to use the software properly; and, some of this is due to getting our arms around the optimal configuration and operation of these packages.

The January 2011 availability level - our lowest month of availability- was due largely to two problems.  One was that we scheduled a maintenance window in our server room so that University of Michigan electricians could wire up a new "whole room" uninterruptible power supply, and this, of course, took our production web systems off-line.  The other problem was that our regular synchronization process between our production systems and our cloud-based replica had failed in an unusual way that was difficult to detect at first.  The database export/import had failed, but only partially, and that produced very odd behavior with our web portal.  It took a significant amount of time to isolate the problem, and by the time we had a workaround deployed, the electricians had finished their work, and the production systems were back on-line.

Monday, June 27, 2011

Watching the Intercloud drift by

I really like this new graphic from DuraCloud
The term Intercloud doesn't get used as often as the term Internet does today, but they may change in a few more years.

Just as the Internet was a "network of networks" the Intercloud is supposed to be a "cloud of clouds."  But is that supposed to mean?

One possible "cloud of clouds" is what DuraSpace is doing with their DuraCloud service.  ICPSR was a pilot tester of DuraCloud, and we will soon sign up as a customer for the newly available DuraCloud production service.

As the nice graphic from DuraCloud makes clear, moving a document into the DuraCloud "cloud" really places it into a collection of other "clouds" as well, making DuraCloud a "cloud of clouds."  From the point of view of a customer like ICPSR, we view DuraCloud as a single location for content with a single interface, a single bill, and a single help desk.  But behind the scenes DuraCloud makes use of other clouds for its storage, such as Amazon's Simple Storage Service (S3) and Rackspace.

And, it is also easy to imagine future "cloud providers" sitting behind DuraCloud where the cloud provider is itself a "cloud of clouds."  For example, the folks behind Chronopolis, a "cloud of clouds" itself with storage locations at the San Diego Supercomputer Center, the National Center for Atmospheric Research, and the University of Maryland's Institute for Advanced Computer Studies, have announced their intention to be one of the storage providers behind DuraCloud.  And so by putting our content into DuraCloud, we may one day also be putting it into Chronopolis, which in turn means putting a copy into the storage clouds at SDSC, NCAR, and UMIACS.