News and commentary about new technology-related projects under development at ICPSR
Friday, February 3, 2012
TRAC: A4.3: Transparency
A4.3 Repository’s financial practices and procedures are transparent, compliant with relevant accounting standards and practices, and audited by third parties in accordance with territorial legal requirements.
The repository cannot just claim transparency, it must show that it adjusts its business practices to keep them transparent, compliant, and auditable. Confidentiality requirements may prohibit making information about the repository’s finances public, but the repository should be able to demonstrate that it is as transparent as it needs to be and can be within the scope of its community.
Evidence: Demonstrated dissemination requirements for business planning and practices; citations to and/or examples of accounting and audit requirements, standards, and practice; evidence of financial audits already taking place.
ICPSR publishes an annual report which contains a financial statement, and it also completes many, many regular reports because it is part of the Institute for Social Research at the University of Michigan.
Wednesday, February 1, 2012
January 2012 web availability
Ick.
If you click on the image, it'll expand into something more easily read. But don't.
January 2012 was not a good month for ICPSR's web delivery system, We missed our monthly goal of 99%+ availability. Not by a lot, but still missed it.
Many of the problems this month are related to an upgrade we made to our delivery infrastructure. We replaced a five-year-old 32-bit machine with a new 64-bit machine with significantly more memory and processing power. And we upgraded from an older version of tomcat to the latest release.
The maintenance itself accounted for less than an hour of downtime, but various issues related to the upgrade brought additional short periods of downtime. A "death by a thousand cuts" sort of thing.
The good news if that the reliability and stability of the system now seems better, and the upgrade has exposed a couple of software flaws that we have been able to correct. It also has eliminated the last of the 32-bit systems from ICPSR, and so we are less likely to run into problems where we (accidentally) try to run 64-bit binaries and libraries on 32-bit machines.
If you click on the image, it'll expand into something more easily read. But don't.
January 2012 was not a good month for ICPSR's web delivery system, We missed our monthly goal of 99%+ availability. Not by a lot, but still missed it.
Many of the problems this month are related to an upgrade we made to our delivery infrastructure. We replaced a five-year-old 32-bit machine with a new 64-bit machine with significantly more memory and processing power. And we upgraded from an older version of tomcat to the latest release.
The maintenance itself accounted for less than an hour of downtime, but various issues related to the upgrade brought additional short periods of downtime. A "death by a thousand cuts" sort of thing.
The good news if that the reliability and stability of the system now seems better, and the upgrade has exposed a couple of software flaws that we have been able to correct. It also has eliminated the last of the 32-bit systems from ICPSR, and so we are less likely to run into problems where we (accidentally) try to run 64-bit binaries and libraries on 32-bit machines.
Monday, January 30, 2012
Customer service, Zingerman's style
Our parent organization, the University of Michigan's Institute for Social Research (ISR), is working with the training component of the Zingerman's family of companies - ZingTrain - to build a customized training module for use at the ISR. The focus is, of course, on delivering excellent customer service, and I had the opportunity to attend a session led by two ZingTrain consultants.
I don't want to give away too much of their "secret sauce" but I found their interaction with the group engaging and informative. I almost used the word "presentation" but that feels wrong; it really isn't a monologue whatsoever. And there are no Powerpoint slides in sight. As you might expect the ZingTrain folks shared some tips and techniques about how they build the right culture and right processes. And they brought goodies from the Bakehouse!
I started to think about some of the tips and techniques I've learned to use in the technology business over the years. In this realm an awful lot of the interaction with others takes place electronically, and so one doesn't have all of the visual cues and tonal cues one normally can use in conversation. For example, how do you let someone know that if the solution you have offered does not work, you want and expect the person to let you know so that you can keep trying to solve the problem? How do you let them know that you will own the problem until it is solved?
One easy way, of course, is to be explicit.
On the other hand, I will often see people write this instead:
And that's my customer service tip for the month.
Hope it helps. :-)
I don't want to give away too much of their "secret sauce" but I found their interaction with the group engaging and informative. I almost used the word "presentation" but that feels wrong; it really isn't a monologue whatsoever. And there are no Powerpoint slides in sight. As you might expect the ZingTrain folks shared some tips and techniques about how they build the right culture and right processes. And they brought goodies from the Bakehouse!
I started to think about some of the tips and techniques I've learned to use in the technology business over the years. In this realm an awful lot of the interaction with others takes place electronically, and so one doesn't have all of the visual cues and tonal cues one normally can use in conversation. For example, how do you let someone know that if the solution you have offered does not work, you want and expect the person to let you know so that you can keep trying to solve the problem? How do you let them know that you will own the problem until it is solved?
One easy way, of course, is to be explicit.
If that doesn't do the trick, please let me know. I have a few other ideas we can try.By asking the person to return and letting them know that "we" can try some other things, it shows that one is engaged. It lets them know that this is the start of a conversation, not the end of one.
On the other hand, I will often see people write this instead:
Hope this helps.I know people often write this with the best of intentions, but consider how people may read it. It sounds like the conversation is over. "Here, try this. I hope it works. But if it doesn't, it's your problem, not mine." There's no invitation to come back for more advice, more assistance, more analysis if the issue hasn't been resolved.
And that's my customer service tip for the month.
Hope it helps. :-)
Friday, January 27, 2012
TRAC: A4.2: Charting and changing a course
A4.2 Repository has in place processes to review and adjust business plans at least annually.
The repository must demonstrate its commitment to proactive business planning by performing cyclical planning processes at least yearly. The repository should be able to demonstrate its responsiveness to audit results, for example.
Evidence: Business plans, audit planning (e.g., scope, schedule, process, and requirements) and results; financial forecasts; recent audits and evidence of impact on repository operating procedures.
It certainly is the case that ICPSR undertakes a great deal of what I will describe as "tactical" business planning, and does so on a regular basis. For example, this is the time of the year when ICPSR begins building its draft budgets for its next fiscal year (July 1 - June 30). For my team this means that we need to look into our crystal balls and guess what our technology equipment and software expenses might be for the following year (7 to 19 months hence) and also allocate our team across two dozen different projects (e.g., 20% of Bob will be on Project X, 10% of Project Y, and the rest of Project Z). My experience is that we almost always create a budget which accurately reflects the needs and direction of the organization, and..... is almost always missing some major initiative that only appears much, much later in the calendar year. And so in practice the technology team aligns about 80-90% of its resources with the plan in the budget, and 10-20% end up doing something unplanned and unbudgeted. This keeps things exciting.
Many of the other workgroups at ICPSR are much smaller than the technology team, and they have relatively long-lived contracts and grants where the budget, deliverables, work scope, etc are reasonably well defined. I'm sure they also encounter their fair share of curveballs from their funding sources, and also review and tweak their project-year budgets at the tactical level.
So maybe a B+ or even an A- overall for ICPSR on "regular, tactical business planning." Not bad.
Many organizations do a much poorer job of looking more deeply into their crystal balls,attempting to peer three, four, maybe even five years down the road. And ICPSR is no exception. I call this "strategic" business planning, and I view its role as complementary to "tactical" business planning.
If "tactical" business planning helps you figure out what you're going to do the next year, and how you're going to get it done, "strategic" business planning helps you figure out what you're NOT going to do any time soon, and how you've going to avoid heading down the wrong roads.
The output of this type of planning isn't a spreadsheet or a list of tasks. It isn't necessarily even a list of goals. Instead it is a shared vision for who you are, what business you are in, and where you think you need to be in three, four, maybe even five years down the road. Here's an example:
Today we're the largest archive of survey research data in the world. Our operations are geared to finding, collecting, curating, preserving, and delivering survey data. We think we are the best at these activities.
In four years, however, we believe that survey data will account for only a tiny portion of our business. We believe that video and social media content will be the new core research content of the future, and this content requires expertise and systems very different than we have today. Our intent is to limit the amount of time we spend growing our staff and growing our systems to support survey data; we will maintain, but not enhance. We will seek out grants and contracts that allow us to build infrastructure and expertise in these areas. And we will begin to invest in our people and our systems for video and social media data.This may or may not be the right vision and the right "strategic" business plan to make, but it illustrates the idea that the organization has charted a course. It lets people know where they are heading. It does NOT tell them how they will get there. (That needs to happen eventually too, of course.) It tells people what they are NOT going to do.
It can be really tough to step outside of the fray and the demands of the day-to-day job to think about the longer term, but it's crucial. Otherwise an organization just keeps heading down the same road instead of looking at other available roads that may lead to better places.
Monday, January 23, 2012
Tech@ICPSR talks about the cloud @ LA2M
Tech@ICPSR will be giving a talk on cloud computing at the February 1, 2011 LA2M meeting. We'll be talking about the cloud; kind of a high-level overview of what different folks say the cloud is, and some of the consumer- and business-oriented services and systems that live in it.
I'll add a link to the materials shortly after the talk, and, if LA2M adds the video of the talk to their archive, I'll add a link to that as well.
I'll add a link to the materials shortly after the talk, and, if LA2M adds the video of the talk to their archive, I'll add a link to that as well.
Friday, January 20, 2012
TRAC: A4.1: Business planning
A4.1 Repository has short- and long-term business planning processes in place to sustain
the repository over time.
The repository must demonstrate that it has formal, cyclical, proactive business planning processes in place. A brief description of the repository’s business plan should show how the repository will generate income and assets through services, third-party partnerships, grants, and so forth. As for A1.2 (succession/ contingency/escrow planning), the repository must establish these processes when it is viable to avoid business crises. These questions may be pertinent to this requirement:
An awful lot of this TRAC requirement is met simply by being a unit of the the Institute for Social Research at the University of Michigan. (This will be a recurring theme across a lot of TRAC A4.) The U-M, like most large universities, has an elaborate bureaucracy for managing budgets and expenditures, and producing a cornucopia of reports, statements, and other fine documents.
In addition, like a lot of non-profits across the US, ICPSR produces an annual report that it makes available publicly, and this contains high-level documentation for both planning and reporting. For example, the ICPSR annual report includes a breakdown of revenues and expenses across major areas, and usually contains several essays on recent accomplishments, upcoming initiatives, and areas of focus for the upcoming fiscal year.
The repository must demonstrate that it has formal, cyclical, proactive business planning processes in place. A brief description of the repository’s business plan should show how the repository will generate income and assets through services, third-party partnerships, grants, and so forth. As for A1.2 (succession/ contingency/escrow planning), the repository must establish these processes when it is viable to avoid business crises. These questions may be pertinent to this requirement:
- Under this plan, to what extent is the repository supported, or expected to be supported, by revenue from content-contributing organizations and agencies, such as publishers?
- To what extent is the repository supported, or expected to be supported, by revenue from subscribers or subscribing institutions?
- What measures are in place, if any, to limit access by nonsubscribing stakeholders?
- What financial incentives are offered, if any, to discourage subscribers from postponing their investment in the repository? From discontinuing investing in the repository?
- To what extent is the repository supported, or expected to be supported, by other kinds of parties?
- How will major future costs, such as migrations, capital improvements, enhancements, providing access in the event of publisher failure, etc., be distributed between publishers, subscribers, and other supporting parties?
- What contingency plans are in place to cover the loss of future revenue and/or outside funding?
- In the event of a catastrophic failure, are reserve assets sufficient to ensure the restoration of subscriber access to content reasonably quickly?
- If this is a national or government-sponsored repository, how is it insulated from political events, such as international conflicts or diplomatic crises, that might affect its ability to serve foreign constituencies?
An awful lot of this TRAC requirement is met simply by being a unit of the the Institute for Social Research at the University of Michigan. (This will be a recurring theme across a lot of TRAC A4.) The U-M, like most large universities, has an elaborate bureaucracy for managing budgets and expenditures, and producing a cornucopia of reports, statements, and other fine documents.
In addition, like a lot of non-profits across the US, ICPSR produces an annual report that it makes available publicly, and this contains high-level documentation for both planning and reporting. For example, the ICPSR annual report includes a breakdown of revenues and expenses across major areas, and usually contains several essays on recent accomplishments, upcoming initiatives, and areas of focus for the upcoming fiscal year.
Wednesday, January 18, 2012
Disaster Recovery v. High Availability
A question I often receive from customers and colleagues is: If ICPSR has a replica of its production delivery system in Amazon's cloud, why is it that the web site is sometimes down due to scheduled maintenance or unplanned outages?
The short answer is: ICPSR's cloud replica serves a disaster recovery (DR) purpose, but not a high availability (HA) purpose. Of course, more often than not, this generates a look that falls somewhere between Bah! and This sounds like some made-up IT nonsense! However, it really is the answer. But that begs the question: What's the difference between DR and HA? But first a trip back in time....
As some long-time ICPSR clients may recall, the ICPSR delivery system was off-line for nearly a week during the holiday break between 2008 and 2009. The root cause was a long power outage due to a major ice storm in the Midwest which knocked out power to many homes and businesses, including many in Ann Arbor. And because ICPSR resides in a building just a little bit off the University of Michigan's central campus, we're just like any other home or business that waits for DTE Energy to restore power.
As one might expect both myself and the ICPSR Director at the time, Myron Gutmann, were quite anxious for the power to be restored. The storm had caused so much damage that it wasn't at all clear when the building's power would be restored. And, after the first few days without power - and heat - the building's pipes were in danger of bursting. Things were looking pretty bad.
However, as it turned out we had been experimenting with Amazon's new computing and storage cloud just prior to the storm. It would be pretty easy to stand up a minimal web server in Amazon's cloud, something that would basically say Yes, we know our delivery system is down, and we're sorry about that. And here's the best guess from the local power company about when power will be restored. We then worked with some of our colleagues at the University of Michigan and the San Diego Computing Center to update the system that maps names (like www.icpsr.umich.edu) to network addresses so that ICPSR's URLs for its web site would point to this new, minimal web server in Amazon's cloud. That didn't fix the problem, of course, but it let people know that ICPSR knew there was a problem, and shared the best information we had about the problem.
Once power was restored and the main delivery system came back on-line, I had a long conversation with Myron about how we wanted to position ICPSR for any future problem like this. What if the building lost power again for an extended period? What if a tornado knocked down the whole building? What if the U-M suffered some catastrophic problem with its network?
One option was to change the architecture of ICPSR's delivery systems. Rather than having a complex series of simple web applications, we could redesign and rebuild the whole system so that it would also contain a middle layer of technology that would catch and route incoming requests to one of many delivery system components. And rather than having a single production system at the University of Michigan, we would build a multi-site production system spread across multiple network providers and service providers so that no single problem would disrupt services. This is essentially the high availability (HA) version of ICPSR's delivery system. It would have the virtue of providing true 99.99%+ reliability, but would cost plenty of money to design, build, and operate. If you are running IT systems for a bank or a hospital or an aircraft carrier, you build them with HA. But what about a data archive?
Another option was to keep the ICPSR delivery architecture the same, but replicate it somewhere off-site. Automated jobs could keep the web content, data content, and web applications synchronized. And an easy - but manual - process could be used to redirect traffic to the replica when needed. In this world there would still be plenty of times where a component of ICPSR's delivery system might be off-line due to maintenance or a fault, but if the maintenance or fault was long-lived, then the replica could be pressed into service. This type of solution would be inexpensive to design, deploy, and operate, and would deliver a credible disaster recovery (DR) story, but would probably only give us uptime somewhere between 99.0% and 99.9%. Would that be good enough?
In the end, of course, we decided that the best use of resources would be to build a system that would still have some outages from time to time, but which would never again be off-line for an entire week. We set an availability goal of 99.5% for each month across all components. That is, every time a single component faults - search, download, online analysis, and so on - it counts against the uptime of the WHOLE system. And we would leave it up to the judgement of the on-call engineer to decide when a problem was likely to be long-lived enough to warrant a switch to the replica.
So we chose DR instead of HA.
Looking back, my sense is that we made the right decision. In practice we seem to hit our 99.5% availability goal most months, and because we did not tie up our software and systems development resources on rebuilding the delivery system to guarantee HA, we were able to design and build systems like our Restricted Contract System, Secure Data Environment, and Virtual Data Enclave. Of course, when we need to perform a major bit of maintenance like last weekend where it is important that we continue to point www.icpsr.umich.edu at the production system rather than the replica, it always makes me wonder about the HA alternative.
The short answer is: ICPSR's cloud replica serves a disaster recovery (DR) purpose, but not a high availability (HA) purpose. Of course, more often than not, this generates a look that falls somewhere between Bah! and This sounds like some made-up IT nonsense! However, it really is the answer. But that begs the question: What's the difference between DR and HA? But first a trip back in time....
As some long-time ICPSR clients may recall, the ICPSR delivery system was off-line for nearly a week during the holiday break between 2008 and 2009. The root cause was a long power outage due to a major ice storm in the Midwest which knocked out power to many homes and businesses, including many in Ann Arbor. And because ICPSR resides in a building just a little bit off the University of Michigan's central campus, we're just like any other home or business that waits for DTE Energy to restore power.
As one might expect both myself and the ICPSR Director at the time, Myron Gutmann, were quite anxious for the power to be restored. The storm had caused so much damage that it wasn't at all clear when the building's power would be restored. And, after the first few days without power - and heat - the building's pipes were in danger of bursting. Things were looking pretty bad.
However, as it turned out we had been experimenting with Amazon's new computing and storage cloud just prior to the storm. It would be pretty easy to stand up a minimal web server in Amazon's cloud, something that would basically say Yes, we know our delivery system is down, and we're sorry about that. And here's the best guess from the local power company about when power will be restored. We then worked with some of our colleagues at the University of Michigan and the San Diego Computing Center to update the system that maps names (like www.icpsr.umich.edu) to network addresses so that ICPSR's URLs for its web site would point to this new, minimal web server in Amazon's cloud. That didn't fix the problem, of course, but it let people know that ICPSR knew there was a problem, and shared the best information we had about the problem.
Once power was restored and the main delivery system came back on-line, I had a long conversation with Myron about how we wanted to position ICPSR for any future problem like this. What if the building lost power again for an extended period? What if a tornado knocked down the whole building? What if the U-M suffered some catastrophic problem with its network?
One option was to change the architecture of ICPSR's delivery systems. Rather than having a complex series of simple web applications, we could redesign and rebuild the whole system so that it would also contain a middle layer of technology that would catch and route incoming requests to one of many delivery system components. And rather than having a single production system at the University of Michigan, we would build a multi-site production system spread across multiple network providers and service providers so that no single problem would disrupt services. This is essentially the high availability (HA) version of ICPSR's delivery system. It would have the virtue of providing true 99.99%+ reliability, but would cost plenty of money to design, build, and operate. If you are running IT systems for a bank or a hospital or an aircraft carrier, you build them with HA. But what about a data archive?
Another option was to keep the ICPSR delivery architecture the same, but replicate it somewhere off-site. Automated jobs could keep the web content, data content, and web applications synchronized. And an easy - but manual - process could be used to redirect traffic to the replica when needed. In this world there would still be plenty of times where a component of ICPSR's delivery system might be off-line due to maintenance or a fault, but if the maintenance or fault was long-lived, then the replica could be pressed into service. This type of solution would be inexpensive to design, deploy, and operate, and would deliver a credible disaster recovery (DR) story, but would probably only give us uptime somewhere between 99.0% and 99.9%. Would that be good enough?
In the end, of course, we decided that the best use of resources would be to build a system that would still have some outages from time to time, but which would never again be off-line for an entire week. We set an availability goal of 99.5% for each month across all components. That is, every time a single component faults - search, download, online analysis, and so on - it counts against the uptime of the WHOLE system. And we would leave it up to the judgement of the on-call engineer to decide when a problem was likely to be long-lived enough to warrant a switch to the replica.
So we chose DR instead of HA.
Looking back, my sense is that we made the right decision. In practice we seem to hit our 99.5% availability goal most months, and because we did not tie up our software and systems development resources on rebuilding the delivery system to guarantee HA, we were able to design and build systems like our Restricted Contract System, Secure Data Environment, and Virtual Data Enclave. Of course, when we need to perform a major bit of maintenance like last weekend where it is important that we continue to point www.icpsr.umich.edu at the production system rather than the replica, it always makes me wonder about the HA alternative.
Subscribe to:
Posts (Atom)


