News and commentary about new technology-related projects under development at ICPSR
Showing posts with label duracloud. Show all posts
Showing posts with label duracloud. Show all posts
Monday, September 24, 2012
DuraSpace announces SDSC as storage partner
DuraSpace announced recently their relationship with SDSC as a storage provider for DuraCloud. As I posted a while ago, we have been using both DuraCloud and their SDSC storage partner for a while. It's great to see DuraCloud continue to grow.
Friday, July 20, 2012
Amazon's loss is SDSC's gain
One of the recent Amazon Web Services (AWS) power outages has left some of my EBS volumes in an inconsistent state. If these were simple volumes, each containing a filesystem, then the fix is easy: just dismount the filesystem, run fsck to check it, and then remount the filesystem after it has been fixed. We have done this on several of our EC2 instances that had inconsistent volumes.
Unfortunately, for these particular volumes we have bonded them together to form a virtual RAID. And this RAID is used as a single multi-TB filesystem which is much bigger than fsck can handle. So we are kind of stuck.
One option would be to newfs the big filesystem, and to move the several TBs of content back into AWS, but that would be very slow. And if there is another power outage......
So instead we called up our pals at Duracloud and asked them if they could help us enable replication of our content to a second provider. (The first provider is - ironically - AWS. But their S3 service, not their EC2/EBS service.) They said they'd be happy to help, and, in fact, they will starting to replicate our content later this same week. (Now that's service!)
The new copy of our content will now be replicated in...... SDSC's storage cloud. This really brings us full circle at ICPSR since our very first off-site archival copy was stored at SDSC. Back then (like in 2008) it was stored in their Storage Resource Broker (SRB) system, and we used a set of command-line utilities to sync content between ICPSR and SDSC.
The SRB stuff was kind of clunky for us, especially given our large number of files, our sometimes large files (>2GB), and our sometimes poorly named files (e.g., control characters in file names). Our content then moved into Chronopolis from SRB, and then at the end of the demonstration project, we asked SDSC to dispose of the copy they had. But now it is coming back......
Unfortunately, for these particular volumes we have bonded them together to form a virtual RAID. And this RAID is used as a single multi-TB filesystem which is much bigger than fsck can handle. So we are kind of stuck.
One option would be to newfs the big filesystem, and to move the several TBs of content back into AWS, but that would be very slow. And if there is another power outage......
So instead we called up our pals at Duracloud and asked them if they could help us enable replication of our content to a second provider. (The first provider is - ironically - AWS. But their S3 service, not their EC2/EBS service.) They said they'd be happy to help, and, in fact, they will starting to replicate our content later this same week. (Now that's service!)
The new copy of our content will now be replicated in...... SDSC's storage cloud. This really brings us full circle at ICPSR since our very first off-site archival copy was stored at SDSC. Back then (like in 2008) it was stored in their Storage Resource Broker (SRB) system, and we used a set of command-line utilities to sync content between ICPSR and SDSC.
The SRB stuff was kind of clunky for us, especially given our large number of files, our sometimes large files (>2GB), and our sometimes poorly named files (e.g., control characters in file names). Our content then moved into Chronopolis from SRB, and then at the end of the demonstration project, we asked SDSC to dispose of the copy they had. But now it is coming back......
Wednesday, November 16, 2011
DuraCloud Archiving and Preservation Webinar
Shameless self-promotion alert...
The nice folks at DuraSpace have published the audio and video from the recent webinar that Michele Kimpton (CEO, DuraSpace) and I gave on DuraCloud.
Michele spends the first 5-10 minutes talking about the business case behind DuraCloud, and then I spend about 30 minutes talking about ICPSR and how we came to use DuraCloud to store a copy of our archival holdings.
The nice folks at DuraSpace have published the audio and video from the recent webinar that Michele Kimpton (CEO, DuraSpace) and I gave on DuraCloud.
Michele spends the first 5-10 minutes talking about the business case behind DuraCloud, and then I spend about 30 minutes talking about ICPSR and how we came to use DuraCloud to store a copy of our archival holdings.
Monday, November 14, 2011
A dangerous combination
Do you like irony?
It turns out that the University of Michigan, like many other organizations, has decided to use the cloud for keeping track of its "Travel and Expense" software and reporting, and has therefore adopted Concur.
I think that the University has made a good decision to put this in the cloud, and to look to use a hosted solution (Software as a Service (SaaS)). Using an existing service makes much more sense than building our own software. How could the U-M build a better application than a company that makes its living doing exactly this sort of thing?
Now, this isn't to say that I am a huge fan of Concur (or at least how it has been implemented at U-M). I don't find the workflow or interface to be all that intuitive, and there are a couple of things that really trip me up all the time. For example, when entering the name of someone, sometimes I am supposed to enter their LAST name and sometimes I am supposed to enter their FIRST name, and I can never remember which to enter. (Cue sad music.)
But the really challenging part about using this cloud service is when I use it to pay for cloud services. (Cue ironic music.)
Each month I get a bill from Amazon. And DuraCloud. And another one from DuraCloud (because we use more space than our membership allows.) And another one from Amazon. (Two different projects with different credit cards and different pools of machines.) And Salesforce. And....
So each month I print the invoice to PDF. And I fetch the receipt from my university credit card, and PDF that too. And then I bundle them together in an expense report in Concur. And that's when the trouble starts: How do I classify the expense?
This is almost certainly not the fault of the Concur software, of course. The problem is in the controlled vocabulary of "expense types" that the U-M has plugged into the system. Not one is a good fit for paying cloud providers. And so I pick one from the choices I do have.
Computer maintenance?
Computer rental?
Memberships (especially for the DuraSpace one, which is indeed a membership)?
Other?
My expense report is reviewed by at least four different people (two within ICPSR, at least one within our parent organization, the Institute for Social Research, at at least one at the U-M central Business and Finance unit). If any one of them believes that I have selected the wrong expense type, the report returns to me, and I then must resubmit it. The good news is that I don't have to reload the invoice or receipt, and so the process is relatively simple.
But for those of you about to implement Concur or another expense and travel reporting system, please add a new expense category for your IT managers: Cloud computing services.
It turns out that the University of Michigan, like many other organizations, has decided to use the cloud for keeping track of its "Travel and Expense" software and reporting, and has therefore adopted Concur.
I think that the University has made a good decision to put this in the cloud, and to look to use a hosted solution (Software as a Service (SaaS)). Using an existing service makes much more sense than building our own software. How could the U-M build a better application than a company that makes its living doing exactly this sort of thing?
Now, this isn't to say that I am a huge fan of Concur (or at least how it has been implemented at U-M). I don't find the workflow or interface to be all that intuitive, and there are a couple of things that really trip me up all the time. For example, when entering the name of someone, sometimes I am supposed to enter their LAST name and sometimes I am supposed to enter their FIRST name, and I can never remember which to enter. (Cue sad music.)
But the really challenging part about using this cloud service is when I use it to pay for cloud services. (Cue ironic music.)
Each month I get a bill from Amazon. And DuraCloud. And another one from DuraCloud (because we use more space than our membership allows.) And another one from Amazon. (Two different projects with different credit cards and different pools of machines.) And Salesforce. And....
So each month I print the invoice to PDF. And I fetch the receipt from my university credit card, and PDF that too. And then I bundle them together in an expense report in Concur. And that's when the trouble starts: How do I classify the expense?
This is almost certainly not the fault of the Concur software, of course. The problem is in the controlled vocabulary of "expense types" that the U-M has plugged into the system. Not one is a good fit for paying cloud providers. And so I pick one from the choices I do have.
Computer maintenance?
Computer rental?
Memberships (especially for the DuraSpace one, which is indeed a membership)?
Other?
My expense report is reviewed by at least four different people (two within ICPSR, at least one within our parent organization, the Institute for Social Research, at at least one at the U-M central Business and Finance unit). If any one of them believes that I have selected the wrong expense type, the report returns to me, and I then must resubmit it. The good news is that I don't have to reload the invoice or receipt, and so the process is relatively simple.
But for those of you about to implement Concur or another expense and travel reporting system, please add a new expense category for your IT managers: Cloud computing services.
Labels:
amazon web servivces,
cloud computing,
duracloud,
duraspace,
fun
Wednesday, October 26, 2011
Using DuraCloud for Archiving and Preservation
I'll be joining Michele Kimpton, CEO of DuraSpace, on a webinar next Wednesday (November 2, 2011). Our topic is DuraCloud, and how one can use this cloud-based service as part of one's digital preservation strategy.
I think sometimes people will view the cloud as an alternative to keeping and maintaining local copies, but at ICPSR we're using the cloud as an easy-to-manage storage location to supplement more conventional locations, such as local NAS storage and the University of Michigan's "Value Storage" service.
Here is a copy of the invite that went out via email:
I think sometimes people will view the cloud as an alternative to keeping and maintaining local copies, but at ICPSR we're using the cloud as an easy-to-manage storage location to supplement more conventional locations, such as local NAS storage and the University of Michigan's "Value Storage" service.
Here is a copy of the invite that went out via email:
| |
|
Monday, September 12, 2011
DuraCloud Pilot talk at Educause 2011
Our colleague from DuraSpace, CEO Michele Kimpton, is giving a talk at this year's Educause conference. Here is a link to the abstract.
The topic is the DuraCloud pilot that concluded earlier this year, and in which ICPSR was an active participant. I've been impressed with the level of commitment and the service orientation of the entire DuraCloud team, and we've since moved from "pilot user" to "customer."
The topic is the DuraCloud pilot that concluded earlier this year, and in which ICPSR was an active participant. I've been impressed with the level of commitment and the service orientation of the entire DuraCloud team, and we've since moved from "pilot user" to "customer."
Monday, June 27, 2011
Watching the Intercloud drift by
![]() |
| I really like this new graphic from DuraCloud |
Just as the Internet was a "network of networks" the Intercloud is supposed to be a "cloud of clouds." But is that supposed to mean?
One possible "cloud of clouds" is what DuraSpace is doing with their DuraCloud service. ICPSR was a pilot tester of DuraCloud, and we will soon sign up as a customer for the newly available DuraCloud production service.
As the nice graphic from DuraCloud makes clear, moving a document into the DuraCloud "cloud" really places it into a collection of other "clouds" as well, making DuraCloud a "cloud of clouds." From the point of view of a customer like ICPSR, we view DuraCloud as a single location for content with a single interface, a single bill, and a single help desk. But behind the scenes DuraCloud makes use of other clouds for its storage, such as Amazon's Simple Storage Service (S3) and Rackspace.
And, it is also easy to imagine future "cloud providers" sitting behind DuraCloud where the cloud provider is itself a "cloud of clouds." For example, the folks behind Chronopolis, a "cloud of clouds" itself with storage locations at the San Diego Supercomputer Center, the National Center for Atmospheric Research, and the University of Maryland's Institute for Advanced Computer Studies, have announced their intention to be one of the storage providers behind DuraCloud. And so by putting our content into DuraCloud, we may one day also be putting it into Chronopolis, which in turn means putting a copy into the storage clouds at SDSC, NCAR, and UMIACS.
Tuesday, May 10, 2011
ICPSR has been busy chatting with our good friends at DuraSpace over the past month or two. We have been an active member of their DuraCloud pilot. This is a hosted service for content and services where the big cloud providers deliver the compute and the storage, and DuraSpace delivers the software and services. Our main use has been as a supplement to our archival storage solution.
The project looks like it will soon finish the jump from "research pilot" to "production service." We participated in a webinar yesterday which introduced the updated management console. It looks good, but we did volunteer one feature request: an "financial administrator" role. The idea is that a login assigned to this role would have read access to invoices and financial statements, but not have any access to the content, services, etc. This is a role we would love to have with the Amazon AWS IAM stuff, but the Amazon guys still haven't identified the monthly bill as one of the system elements that would benefit from such a role. (And so that means that someone like me has to navigate through the management console to grab billing information each month, then save it, upload it into the absolutely ghastly UMich financial systems, and ....)
DuraCloud is a nice fit for ICPSR since it gives us a single management interface for syncing content to multiple cloud providers (Amazon and Rackspace today, but Microsoft down the road), and for invoking preservation-oriented services over the content, such as fixity checking.
You can find more info about DuraCloud on their web site, and a nice little piece they wrote about ICPSR too.
Wednesday, December 8, 2010
DuraCloud pilot update - December 2010
I'm getting together with my DuraCloud pilot colleagues the morning before CNI starts. I just saw the agenda for the meeting, and it looks like it will be a fruitful and interesting set of conversations.
I also had a very nice chat with Carol Minton Morris from Duraspace about our expectations from the DuraCloud pilot project, and our experience to date. You can find her write-up of that conversation here on the DuraSpace blog.
I also had a very nice chat with Carol Minton Morris from Duraspace about our expectations from the DuraCloud pilot project, and our experience to date. You can find her write-up of that conversation here on the DuraSpace blog.
Labels:
cloud computing,
cyberinfrastructure,
duracloud,
duraspace
Tuesday, November 23, 2010
The Cloud and Archival Storage
Price. Availability. Services. Security.
These are the four parameters that I use when deciding where to store one of our archival storage copies.
For me the cloud is just another storage container. Fundamentally it is no different from a physical storage location except in how it differs across these four dimensions. In fact, I can conceptualize my "non-cloud" storage locations as storage as a service cloud providers, but where the provider is a lot more local than the big names in "cloud" today:
ICPSR Cloud: This is the portion of the EMC NS-120 NAS that I use for a local copy of archival storage. It is very expensive with a reasonably high-level of availability. It provides very few services; if I want to perform a fixity check of the objects I have stored here, I have to create and schedule that myself. Because I have physical control over the ICPSR Cloud, I have an irrational belief that it is probably secure, even though I know that ICPSR isn't as physically secure as many other companies at which I have worked. Certainly ICPSR does not make any statements or guarantees about ISO 27001 compliance.
UMich Cloud: This is a multi-TB chunk of NFS file storage that I rent from the University of Michigan's Information Technology Services (ITS) organization. They call it ITS Value Storage. The price here is excellent, but the level of availability is just a hair lower. I don't notice the lower level of availability most of the time, but I do perceive it when running long-lived, I/O-intensive applications. Like my own cloud, this one has no services unless I deploy them myself. Because I do not have physical control over the equipment, or even know exactly where the equipment is (beyond a given data center), it feels like there is less control. ITS makes no promises about ISO 27001 compliance (or promises about other standards), but my sense is that their controls and physical security and IT management processes must be at least as good as mine. After all, they are managing many, many TBs for many different university departments and organizations, including themselves.
Amazon Cloud: This is a multi-TB chunk of Elastic Block Storage (EBS) that I rent from Amazon Web Services. I use EBS rather than the Simple Storage Service (S3) because I want the semantics of a filesystem so that I don't have to worry about things like files that are large or that have funny characters in their names. The price here is good, better than my EMC NAS, but not as good as the ITS Value Storage. The availability is quite good overall, but, of course, the network throughput between ICPSR and AWS is nowhere near as good as intra-campus networking, and it is even worse for the AWS EU location. The services are no better and no worse than my own cloud or the UMich cloud. Like the ITS Value Storage service I have no control over the physical systems, and I know even less about their physical location. Amazon says that it passed a SAS 70 audit, and recently received an ISO 27001 certification. This seems to be a better security story than anyone else so far.
DuraCloud: Unlike the other clouds, I'm not using this one for archival storage; it is still in a pilot phase. The availability is similar to plain old AWS (which hosts the main DuraCloud service), and the price is still under discussion. My expectation is that the level of security is no better (and no worse) than the underlying cloud provider(s), and so depending upon which storage provider one selects, one's mileage may vary. However, the really interesting thing about DuraCloud is the idea of the services. If DuraCloud can execute and deliver useful, robust services on top of basic storage clouds, that will be a true value-add, and will make this a very compelling platform for archival storage.
Chronopolis: Like DuraCloud, this too is not in production yet, and is being groomed (I think) as a future for-fee, production service. I don't have as much visibility here with regard to availability since I am not actively moving content in and out of Chronopolis; most of the action seems to be taking please under the hood between the storage partner locations. My sense is that the level of security is probably similar to the UMich Cloud since the lead organization, the San Diego Supercomputer Center, is in the world of higher education, like UMich, but it may well be the case that they have a stronger security story to tell, and I just don't know it. And like DuraCloud, my sense is that it will come down to services: If Chronopolis builds great services that facilitate archival storage, that will make it an interesting choice.
These are the four parameters that I use when deciding where to store one of our archival storage copies.
For me the cloud is just another storage container. Fundamentally it is no different from a physical storage location except in how it differs across these four dimensions. In fact, I can conceptualize my "non-cloud" storage locations as storage as a service cloud providers, but where the provider is a lot more local than the big names in "cloud" today:
ICPSR Cloud: This is the portion of the EMC NS-120 NAS that I use for a local copy of archival storage. It is very expensive with a reasonably high-level of availability. It provides very few services; if I want to perform a fixity check of the objects I have stored here, I have to create and schedule that myself. Because I have physical control over the ICPSR Cloud, I have an irrational belief that it is probably secure, even though I know that ICPSR isn't as physically secure as many other companies at which I have worked. Certainly ICPSR does not make any statements or guarantees about ISO 27001 compliance.
UMich Cloud: This is a multi-TB chunk of NFS file storage that I rent from the University of Michigan's Information Technology Services (ITS) organization. They call it ITS Value Storage. The price here is excellent, but the level of availability is just a hair lower. I don't notice the lower level of availability most of the time, but I do perceive it when running long-lived, I/O-intensive applications. Like my own cloud, this one has no services unless I deploy them myself. Because I do not have physical control over the equipment, or even know exactly where the equipment is (beyond a given data center), it feels like there is less control. ITS makes no promises about ISO 27001 compliance (or promises about other standards), but my sense is that their controls and physical security and IT management processes must be at least as good as mine. After all, they are managing many, many TBs for many different university departments and organizations, including themselves.
Amazon Cloud: This is a multi-TB chunk of Elastic Block Storage (EBS) that I rent from Amazon Web Services. I use EBS rather than the Simple Storage Service (S3) because I want the semantics of a filesystem so that I don't have to worry about things like files that are large or that have funny characters in their names. The price here is good, better than my EMC NAS, but not as good as the ITS Value Storage. The availability is quite good overall, but, of course, the network throughput between ICPSR and AWS is nowhere near as good as intra-campus networking, and it is even worse for the AWS EU location. The services are no better and no worse than my own cloud or the UMich cloud. Like the ITS Value Storage service I have no control over the physical systems, and I know even less about their physical location. Amazon says that it passed a SAS 70 audit, and recently received an ISO 27001 certification. This seems to be a better security story than anyone else so far.
DuraCloud: Unlike the other clouds, I'm not using this one for archival storage; it is still in a pilot phase. The availability is similar to plain old AWS (which hosts the main DuraCloud service), and the price is still under discussion. My expectation is that the level of security is no better (and no worse) than the underlying cloud provider(s), and so depending upon which storage provider one selects, one's mileage may vary. However, the really interesting thing about DuraCloud is the idea of the services. If DuraCloud can execute and deliver useful, robust services on top of basic storage clouds, that will be a true value-add, and will make this a very compelling platform for archival storage.
Chronopolis: Like DuraCloud, this too is not in production yet, and is being groomed (I think) as a future for-fee, production service. I don't have as much visibility here with regard to availability since I am not actively moving content in and out of Chronopolis; most of the action seems to be taking please under the hood between the storage partner locations. My sense is that the level of security is probably similar to the UMich Cloud since the lead organization, the San Diego Supercomputer Center, is in the world of higher education, like UMich, but it may well be the case that they have a stronger security story to tell, and I just don't know it. And like DuraCloud, my sense is that it will come down to services: If Chronopolis builds great services that facilitate archival storage, that will make it an interesting choice.
Monday, October 25, 2010
DuraCloud pilot update - October 2010
ICPSR's participation in the DuraCloud pilot is coming along nicely. While we were not one of the original pilot members, we were one of the "early adopters" in the second round of pilot users. We've been using DuraSpace pretty actively since the summer.
The collection we selected for pilot purposes is a subset of our archival storage that contains preservation copies of our public-use datasets and documentation. This is big but not too big (1TB or so), and contains a nice mix of format types, such as plain text, XML, PDF, TIFF, and more. At the time of this post, we have 72,134 files copied into DuraCloud.
I've been using their Java-based command-line utility called the synctool to synchronize some of our content with DuraCloud. I found it useful to wrap the utility in a small shell script so that I do not need to specify as many command-line arguments when I invoke it. I tend to use sixteen threads to synchronize content rather than the default three, and while that places a heavy load on our machine here, it leads to faster synchronization. The synctool assumes an interactive user, and has a very basic interface for checking status.
Overall I like the synctool but wish that it had an option that did not assume an interactive user; something I could run out of cron like I often do with rsync. Because the underlying storage platform (S3) limits the size of files, synctool is not able to copy some of our larger files. I wish synctool would "chunk up" the files into more manageable pieces, and sync them for me. One reason I don't use raw S3 for storage is because of this file size limitation; instead I like to spend a little more money and attach an Elastic Block Storage volume (S3-backed) to a running instance, and then use the filesystem to hide the limitation. Then I can just use standard tools, like rsync, to copy very large files into the cloud.
The DuraCloud folks have been great collaborators: extremely responsive, extremely helpful; just a joy to work with. They've told me about a pair of upcoming features that I'm keen to test.
One, their fixity service will be revamped in the 0.7 release. It'll have fewer options and features, but will be much easier to use. I'm eager to see how this compares to a low-tech approach I use for our archival storage: weekly filesystem scans + MD5 calculations compared to values stored in a database.
Two, their replicate-on-demand service is coming, and ICPSR will be the first (I think) test case to replicate its content from S3 to Azure's storage service. I have not had the opportunity to use Microsoft's cloud services at all, and am looking forward to seeing how it performs.
The collection we selected for pilot purposes is a subset of our archival storage that contains preservation copies of our public-use datasets and documentation. This is big but not too big (1TB or so), and contains a nice mix of format types, such as plain text, XML, PDF, TIFF, and more. At the time of this post, we have 72,134 files copied into DuraCloud.
I've been using their Java-based command-line utility called the synctool to synchronize some of our content with DuraCloud. I found it useful to wrap the utility in a small shell script so that I do not need to specify as many command-line arguments when I invoke it. I tend to use sixteen threads to synchronize content rather than the default three, and while that places a heavy load on our machine here, it leads to faster synchronization. The synctool assumes an interactive user, and has a very basic interface for checking status.
Overall I like the synctool but wish that it had an option that did not assume an interactive user; something I could run out of cron like I often do with rsync. Because the underlying storage platform (S3) limits the size of files, synctool is not able to copy some of our larger files. I wish synctool would "chunk up" the files into more manageable pieces, and sync them for me. One reason I don't use raw S3 for storage is because of this file size limitation; instead I like to spend a little more money and attach an Elastic Block Storage volume (S3-backed) to a running instance, and then use the filesystem to hide the limitation. Then I can just use standard tools, like rsync, to copy very large files into the cloud.
The DuraCloud folks have been great collaborators: extremely responsive, extremely helpful; just a joy to work with. They've told me about a pair of upcoming features that I'm keen to test.
One, their fixity service will be revamped in the 0.7 release. It'll have fewer options and features, but will be much easier to use. I'm eager to see how this compares to a low-tech approach I use for our archival storage: weekly filesystem scans + MD5 calculations compared to values stored in a database.
Two, their replicate-on-demand service is coming, and ICPSR will be the first (I think) test case to replicate its content from S3 to Azure's storage service. I have not had the opportunity to use Microsoft's cloud services at all, and am looking forward to seeing how it performs.
Wednesday, September 22, 2010
DuraCloud fixity service testing
Our DuraCloud pilot test is going well. We have uploaded a test collection of nearly 70k files, representing that portion of our archival content that contains public-use datasets. (The datasets are public-use, but our licensing terms restrict access to some of these to our member institutions.)To the left you can see a snapshot from the DurAdmin webapp that one uses to manage content. I've been using this webapp to view content, check progress, and download files. I've been using a command-line utility called synctool for copying content from ICPSR into DuraSpace, and keeping it synchronized.
The image to the left is the right-side panel from the Services tab of the DurAdmin webapp. I've deployed the Fixity service, and am using it to check the bit-level integrity of the content.
I started the service earlier this morning, and it still has quite a bit of work left to do. The processing-status line shows that the service has started, and that it is checked about 4300 of the files so far.
Wednesday, August 25, 2010
DuraCloud pilot update

Things are starting to move along nicely with ICPSR's participation in the DuraCloud pilot. My early experience with the software and tools is mostly positive, but they are still clearly a bit rough around the edges. For example, I've run into bugs on the login screen of the DuraCloud Admin web app that I would characterize as minor, such as needing to use the Submit button on the screen and not the Enter key on the keyboard for some browsers. That said, the DuraSpace people have been fabulous: It's clear they care a lot about the project and the pilot testers, and they have been very, very responsive.
ICPSR is going to test out three parts of DuraCloud.
One, we'll execute a basic upload test, moving content from ICPSR to a "space" in DuraCloud. For this test I'll be using the DuraCloud Admin tool to create a "space," which is basically the same thing as a "bucket" in Amazon's S3. Then I'll use the DuraCloud "sync tool" to copy a subset of ICPSR's archival content to DuraCloud.
Two, we're going to help spec out a "dashboard" or high-level view that shows the integrity of a collection in DuraCloud, and then execute a "fixity" test to measure performance and reliability. For example, if I have a 1TB collection in DuraSpace, and I have that collection replicated across N cloud storage providers, what's my cost to execute such a test every week?
Three, we're going to test a "replication helper" utility that facilitates replicating content across cloud providers. This is a very compelling service for me. Since we already make extensive use of AWS, using DuraCloud as a front-end to AWS is not very interesting for us; but, if we can use DuraCloud as a single front-end to AWS and RackSpace and Atmos and .... then things get more interesting since it means we don't have to develop expertise with ALL of the cloud providers.
Tuesday, June 15, 2010
DuraCloud Pilot
ICPSR is one of several organizations which are participating in "Round #2" of the DuraCloud pilot. This phase of the pilot begins in September and runs through the end of the calendar year.I had missed two earlier webinar sessions on the technology and the pilot process, but read through the slides and listened to the audio today. It really looks like DuraCloud will be an interesting project. In brief it is essentially an abstraction layer on top of multiple cloud storage providers, and also implements a series of services, such as a bit-integrity checker.
For our participation I think we'll use our archival content, but nothing that involves any level of confidentiality. So fair game would be previous (and current) versions of public-use datasets and documentation files (codebooks), and perhaps snapshots of study-level metadata encapsulated in DDI XML.
If the pilot goes well and DuraCloud looks like an attractive service for making additional preservation copies of materials, a future project might be to design a system for encrypting our more sensitive content so that it too could be placed in cloud storage.
Monday, October 5, 2009
DuraCloud

I attended a webinar on DuraSpace last Wednesday. As a big fan of "the cloud" I was very interested to hear about what's been built, how it could be used, and a roadmap of the future. I learned a little bit about all three topics on the webinar.
Gina Jones from the Library of Congress hosted the meeting, and the main speaker was Michele Kimpton.
DuraCloud is being built as an OSGi container sitting on top of cloud storage providers. Customers can view DuraCloud as a buying club for lower prices, and for easing the burden of learning the administrative and software interfaces of each cloud provider.
DuraCloud is starting a pilot project with four cloud providers: (1) Amazon, (2) EMC, (3) Rackspace, and (4) Sun. They are also working actively to add Microsoft as a fifth cloud provider. They have two content providers signed up for the pilot: the New York Public Library, and the Biodiversity Heritage Library.
The NYPL has 800k objects and 50TB of content. They'd like to use DuraSpace to make a copy of their materials, and to transform content from TIFF format to JPEG2000. The JPEG2000 images would then be pulled back out of the cloud to local storage at the NYPL.
The BHL has 40TB of content, and is hoping to use DuraCloud to distribute its content across multiple locations (US, EU), and as a platform for hosting computational intensive data mining.
The pilot is running through the end of the calendar year, and DuraSpace intends to have a pricing model in place by Q2 2010, and to launch a production service in Q3.
In response to a question from a participant, Michele indicated that the focus was NOT on securing sensitive data, but rather on hosting public data with open access. So DuraCloud might be a good bet for some of the content ICPSR delivers on its web site, for example, but not for medical records, confidential data, etc.
Subscribe to:
Posts (Atom)














