Showing posts with label stats. Show all posts
Showing posts with label stats. Show all posts

Wednesday, July 17, 2013

ICPSR Web Availability - 2012-2013

Here are the final numbers for ICPSR's web site availability over our last fiscal year:

Click to embiggen
The year did not start off so well, and we reached the nadir quickly.  August 2012 was our worst period of availability in a very bad year for us overall.  January, March, and June 2012 also had very poor numbers.

The main antagonist we faced was a new and unusual problem with our Oracle database server.  For many years we would export the content for backup purposes each evening, and it worked well for a decade.  However, suddenly in 2012 we began to experience an outage just AFTER each export.  Despite intensive analysis by ourselves and local Oracle exports, we never could isolate the root cause of failure.

We eventually "solved" the problem by exporting our database only once per week v. once per day.  That left us more exposed to loss, of course, but it seemed to limit the outages to once per week v. once per day.

We then replaced the hardware with a new machine with a bit more processor and memory, but with blindingly fast solid-state drives. With the new machine deployed we returned to our daily export schedule, and the machine -- and our web availability -- have been in pretty good shape ever since. The machine went into service in April 2012, and the chart above makes it clear that life has been a little less hectic for our on-call engineer since then.

Thursday, April 18, 2013

Web availability at ICPSR - March 2013

ICPSR's content delivery system showed very high availability in March 2013:  a bit over 99.95% uptime.  We had only two problems in March.  One was a power outage that affected our headquarters on the University of Michigan campus, and we experienced a small amount of downtime as we moved service to our replica in Amazon's cloud.  The second was a 21-minute outage due to a continuing -- but now solved, we think -- problem with exporting content from our Oracle database server.

Here are the overall numbers for ICPSR's 2012-2013 fiscal year:


click to enlarge

We replaced our aging Oracle database server with a new machine which has twice the memory, twice the computing power, and perhaps most impressively, has 300 times the disk I/O speed(!).  The new machine has an array of solid-state drives (SSDs), and we use this for all of our database storage.  (The operating system resides on conventional disk drive technology.)

Friday, November 16, 2012

Web availability at ICPSR - October 2012

October was a very good month for system uptime - over 99.9% availability:

Click chart to enlarge
That's good news after a much rougher September.  So far things look good this month, although a number of very short-lived outages have already pushed us below 99.9% for the month.

Wednesday, October 10, 2012

September 2012 deposits at ICPSR

The numbers from September are in:


# of files# of depositsFile format
291F 0x07 video/h264
14517application/msword
51application/octet-stream
29211application/pdf
145application/vnd.ms-excel
11application/vnd.ms-powerpoint
1201application/x-arcview
311application/x-dbase
11application/x-rar
246application/x-sas
146application/x-spss
135application/x-stata
55application/x-zip
11image/jpeg
321image/x-3ds
702multipart/appledouble
103text/plain; charset=unknown
6218text/plain; charset=us-ascii
21text/rtf
292text/xml

Interesting month in that we have the usual stuff in the usual quantities, but we also have a large number of unusual formats hitting the doorstep, such as ArcView and Apple Double.  And we also have a usual format in an unusually high quantity (MS Word).

Friday, October 5, 2012

September 2012 web availability

September was an OK, but not great month for web availability:


Click to enlarge
We eliminated one frequent, but short-lived source of downtime when we stopped exporting the content of our Oracle database nightly.  We are now doing it only on the weekend, and while that adds some risk, we're gaining significant uptime.  (For some reason that we do not understand, our Oracle instance stops answering queries for 15-20 minutes about ten minutes AFTER the export completes.)  We have a new server racked and ready to install, and we're hoping that a fast new machine with solid-state drives will solve the problem for us.

We did run into some trouble mid-month when some routine maintenance went awry, and we had to fail over to our replica over the weekend of September 15 and 16.  The total amount of downtime was about 90 minutes total over the course of the weekend, but the replica kept the problem from clobbering our service completely.

After that we had pretty smooth sailing for the rest of the month.  Just 16 minutes of downtime for the rest of the month.

Friday, September 7, 2012

August 2012 deposits

Light month for deposits:

# of files# of depositsFile format
5630application/msword
6926application/pdf
98application/vnd.ms-excel
21application/vnd.wordperfect
41application/x-dosexec
247application/x-sas
4721application/x-spss
72application/x-stata
262application/x-zip
11image/gif
184image/jpeg
11message/rfc8220117bit
33text/plain; charset=iso-8859-1
1710text/plain; charset=unknown
329text/plain; charset=us-ascii
32text/plain; charset=utf-8
397text/rtf
11text/x-mail; charset=iso-8859-1
42text/x-mail; charset=unknown
22text/x-mail; charset=us-ascii
11text/xml

Just the usual stuff, but in pretty low quantities.

Wednesday, September 5, 2012

ICPSR web availability - August 2012



August was not our best month.

We did a bit better than 99.3% uptime.  Almost all of the downtime is due to a recurring, as-yet-unsolved problem we are having with our Oracle database platform.  The primary symptom is that the database platform stops fielding queries for about 5-15 minutes, which disables our production web site.  The platform does this about 30-45 minutes AFTER it has finished a full export using the Oracle datapump system.

Because our existing Oracle hardware is old and has a relatively slow disk I/O system, we're going to try to solve this problem by throwing hardware at it.  For well under $10k we can replace our five-year-old hardware with something much newer.  Goodbye RAID-5 SCSI, hello SSD.


Friday, August 10, 2012

July 2012 web availability at ICPSR

July 2012 was a pretty good month:

Clicking the image will open a larger, easier-to-read chart.

We only has 36 minutes of downtime in July, and 28 were due to maintenance as we tried (but failed) to update apache, perl, and mod_perl on our production web server.  We discovered some interesting idiosyncrasies in some perl libraries during the maintenance.  (Summary:  Multi-word time zones like "New York" are trouble.)

Wednesday, August 8, 2012

July 2012 deposits at ICPSR

The totals for July 2012:

# of files# of depositsFile format
11F 0x07 video/h264
52011application/dicom
11application/msaccess
11223application/msword
1214application/octet-stream
24141application/pdf
65application/vnd.ms-excel
11application/vnd.ms-powerpoint
11application/x-7z-compressed
441application/x-arcview
831application/x-dbase
733application/x-dosexec
51application/x-empty
113application/x-sas
11application/x-shellscript
15021application/x-spss
236application/x-stata
284application/x-zip
100245image/jpeg
81image/png
41image/x-ms-bmp
22message/rfc8220117bit
151multipart/appledouble
55text/html
11text/plain; charset=iso-8859-1
43text/plain; charset=unknown
15738text/plain; charset=us-ascii
74text/plain; charset=utf-8
74text/rtf
43text/x-mail; charset=us-ascii
434text/xml
251video/unknown


Lots of image content this month to go with the usual stuff (e.g., SPSS) in the usual volumes.

Wednesday, August 1, 2012

ICPSR 2012 technology recap - where did the money go?

A post from a week or so ago showed the sources of money flowing into the technology organization at ICPSR.  This week's post will focus on how that money is spent.

I should note that the focus of this post is on what we call the "Computer Recharge" at ICPSR.  This is a tax paid by all FTEs at ICPSR that is levied on an hourly basis.  Each time an employee completes his or her timesheet and allocates time to a project (exceptions: sick, holiday, and vacation time), a small amount of money also accumulates in the "Recharge."  In FY 2012 this accounted for about 45% of the technology revenue.

Unlike a direct charge for technology where a project or grant may be paying for a dedicated server, extra storage space, or a fraction of a software developer working on custom systems, the Computer Recharge dollars are used to fund technology expenses that benefit the entire organization.  This includes expenses as prosaic as printers and desktop computers, but also includes systems and software development for our delivery systems, ingest systems, and digital preservation systems.

Here's the breakdown in chart form:

No surprise that most of the money goes to pay people:  systems administrators, desktop support specialists, software developers, and a group of very hands-on managers who do a lot of the same work plus project management and business analysis.  This accounts for $862k out of $1218k.

Equipment ($214k) represents desktop machines, new storage capacity, printers, virtual machines, cloud storage, software licenses, and almost every non-salary expense.

Transfers ($103k) is an interesting category.  This is money that we collect as major systems depreciate.  But since the U-M doesn't really "do" depreciation, we use this interesting process instead:


  1. Buy the item (using money from Equipment pot)
  2. Estimate the lifetime of the item (say, five years)
  3. Collect 1/5 of the purchase price each year for five years
  4. At the end of each year, move the 1/5 collected into Transfers
  5. When the item needs to be replaced, use the money in the Transfers pot


So the $103k represents money that was collected in FY 2012 for items that were purchased in earlier years, and which will need to be replaced in FY 2013 or beyond.

All of our other expenses are tiny by comparison:  $19k for Internet access, $7k for telephones and service monitoring, $6k for travel, $5k for maintenance contracts, and $2k for miscellaneous fees and expenses.

One take-away from this is that the essential element in technology budgeting isn't the purchase price of the server, or the annual cost of cloud storage, bur rather the recurring cost of the people who will build, maintain, enhance, and customize the technology portion of the business.

Monday, July 16, 2012

ICPSR 2012 technology recap - where did the money come from?

We're putting together some summary numbers for technology spending and investments at ICPSR for FY 2012.  (The ICPSR fiscal year is the same as the University of Michigan's, and runs from July 1 to June 30.  We've just recently closed FY 2012.)

The first set of numbers shows the allocation of effort in FY 2012 by funding source. The unit of measurement in this pie chart is HOURS (not DOLLARS) that were expended in FY 2012 by each funding source.  (We originally wanted to calculate dollars, but that turns out to be an even bigger effort.)  Here's an interactive chart:



This is an interactive Google Docs chart.  If you click slices of the pie, it will identify the funding source.

The main source of technology effort funding comes from the Computer Recharge, an hourly "tax" that ICPSR levies against all projects.  Although it is one single funding source (nearly 45% of hours worked in FY 2012 were billed against this source), I have split it into two sub-categories, one for what I am calling "IT" and one for "SW" (software).

The "SW" portion includes the effort of all staff who are professional software developers.  The type of work performed by this team using this account includes enhancements and maintenance for ICPSR's core data curation and data management systems, and investments in new products and services such as software developed to support our IDARS system for applying for access to datasets.

The "IT" portion includes the effort of the remainder of the staff which tends to include systems administrators, architects, network managers, and desktop support specialists.  I also allocate my own time to this bucket since the majority of my non-contract, non-grant effort over the past year has been in building and architecting technology systems.

Other big slices of the "IT pie" include the work of staff members who are explicitly funded by projects such as our CCEERC and RCMD web portals; our two Bill and Melinda Gates Foundation grants; the ICPSR Summer Program, and many more.  In fact, there are over 20 separate funding sources used to support technology at ICPSR; the pie chart shows 18 because I grouped several small ones into a category called "Misc."

If this gives the impression that there are many, many projects and activities at ICPSR that involve technology, that's good!  That is certainly the case.

However, if "focus wins" then we're in a little bit of trouble.  My sense is that each of this 20-some funding sources has at least one unique project with its own business analysis and project management needs, and it is sometimes the case that different projects have antithetical technology needs.  I see this play out in all phases of the OAIS lifecycle.  ("I want you to build a system that makes it as easy as possible to fetch datasets from ICPSR" v.  "I want you to build a system that requires significant effort and oversight to fetch datasets from ICPSR.")


Monday, July 9, 2012

June 2012 deposits at ICPSR

Deposit numbers from June:

# of files# of depositsFile format
11video/h264
36817application/msword
31application/octet-stream
27626application/pdf
94application/vnd.ms-excel
11application/vnd.ms-powerpoint
11application/x-rar
83application/x-sas
11816application/x-spss
7138application/x-stata
11image/jpeg
61image/x-3ds
221multipart/appledouble
44text/html
361text/plain; charset=iso-8859-1
313text/plain; charset=unknown
38910text/plain; charset=us-ascii
11text/plain; charset=utf-8
43text/x-mail; charset=us-ascii
11video/unknown


Quite a bit of Stata this month, much more than normal.

Monday, July 2, 2012

Amazon Web Services makes Tech@ICPSR weep

June 2012 was looking to be a great, great month for uptime.  We were on track to have our best month since November 2011 this fiscal year - just 60 minutes of downtime across all services and all applications.  It was going to be beautiful.

And then Amazon Web Services had another power failure.

And then we wept.

The power failure took the TeachingWithData portal out of action.  (To be fair, it was already having significant problems due to its creaky technology platform, but this took it all the way out of action.)  The failure also took our delivery replica out of action, and gave Tech@ICPSR the joy of rebuilding it over the weekend.

But the real trouble was with a company called Janrain.

Janrain sells a service called Engage.  Engage is what allows content providers (like ICPSR) to use identity providers (like Google, Facebook, Yahoo, and many more) so that their clients (like you) do not need to create yet another account and password.  Engage is a hosted solution that we use for our single sign-on service using existing IDs, and it works 99.9% of the time.

However, this hosted solution lives in the cloud.  We just point the name signin.icpsr.umich.edu at an IP address we get from Janrain, plug in calls to their API, and then magic happens.

Except when the cloud breaks.

Amazon took Engage off-line for nearly four hours.  And then once it came back up, it was thoroughly confused for another three hours.  Ick.

So, counting all of that time as "downtime" our fabulous June 2012 numbers suddenly became our awful June 2012 numbers.  Here they are:





If you click on the image above, Blogger will make it bigger.

Of course, during a lot of that downtime, all of the features on the web site except for third-party login worked fine.  And most of the problem happened late on a Friday night and Saturday morning during the summer, so that's a good time for something bad to happen, if it has to happen at all.

Monday, June 11, 2012

May 2012 deposits at ICPSR

Stats?  Stats.

# of files# of depositsFile format
11application/msaccess
56422application/msword
3343application/octet-stream
13622application/pdf
43application/vnd.ms-excel
161application/x-arcview
41application/x-dbase
317application/x-sas
28119application/x-spss
173application/x-stata
53application/x-zip
187image/jpeg
22message/rfc8220117bit
43text/html
44text/plain; charset=iso-8859-1
1253text/plain; charset=unknown
33220text/plain; charset=us-ascii
11text/plain; charset=utf-8
93text/rtf
11text/x-makefile; charset=us-ascii

Nothing too interesting this month.  We have the usual formats, and in the usual proportions.  We did seem to get an unusually large number of MS Word files last month, and we also have a pretty large set of unidentified files (at least in terms of MIME type).

Friday, June 8, 2012

May 2012 web availability

Web availability was good, but not great, in May 2012:

Click to enlarge


Five main episodes of 25 to 59 minutes account for almost all of the 242 minutes of unavailability in May 2012.

On May 13 the production web server seemingly lost power, and it required an additional reboot and some TLC to bring the system back on-line fully (33 minutes).

On May 18 the search index on our CCEERC web portal became corrupted, and that disabled much of the usefulness of the site for nearly an hour (49 minutes).

On May 22 we saw the first of two episodes where the proxy (AJP) between Apache httpd and Apache tomcat faulted.  This did not recover on its own and required some help from the technology team.  This resulted in a medium-duration outage (27 minutes), and another similar fault occurred on the evening of May 31 (24 minutes).

On May 31 our production database server faulted, requiring a manual power-cycle, and also requiring the production web server to be rebooted (59 minutes).

My sense is that while we're in much better shape with regard to the instability caused by khugepaged, we are starting to see something a little amiss with the Apache proxy system.  It isn't clear to us at this time if the issue is faulty software or sub-optimized configuration on our part.

Friday, June 1, 2012

ICPSR system outage - 5/31/2012

ICPSR's content delivery systems faulted at approximately 8:30pm EDT on Thursday, May 31, 2012.  The oncall engineer discovered that the production Oracle database server had become unresponsive, and this disabled most features of most of our web portals.

After arriving on-site she rebooted the database server, but by then the production web server had become hopelessly confused.  She then rebooted that system as well, and all systems were back in service a bit before 9:30pm EDT.

The ICPSR technology team is reviewing system logs and access records to see if any further corrective action is required.

Our apologies for the inconvenience this no doubt caused to many of you.

Friday, May 4, 2012

April 2012 deposits at ICPSR

Chart?  Chart.

# of files# of depositsFile format
1771906application/dicom
4923application/msword
1576application/octet-stream
18436application/pdf
84application/vnd.ms-excel
11application/x-7z-compressed
17177application/x-dosexec
62application/x-empty
134application/x-sas
2115application/x-spss
33application/x-stata
77application/x-zip
62image/gif
15794813image/jpeg
2086image/png
1046image/x-ms-bmp
66message/rfc8220117bit
52multipart/appledouble
267text/html
22text/plain; charset=unknown
37135text/plain; charset=us-ascii
1046text/plain; charset=utf-8
5912text/rtf
4166text/xml

Things look pretty light in terms of the usual formats, but a handful of big, big deposits stick out.

DICOM - Digital Imaging and Communications in Medicine - is a format used for medical imaging, and we received well over 150k such files last month plus an equally large batch in the more prosaic JPEG format.  I'm not sure how we'll be curating and delivering this content.....

Related to this same set of deposits is a large number of Windows executable files (EXE and DLL file extensions) which will also be an interesting challenge for delivery.

Wednesday, May 2, 2012

April 2012 Web availability

April was a much, much better month for our systems:

Click to enlarge
The real game changer seemed to be disabling the transparent hugepage system on our RHEL 6 systems.  Once we did that, our fortunes changed for the better.  And so we sing:


The main culprits behind the small amount of downtime we had in April were a misfire during an attempt to introduce yet another rewrite rule to our Apache httpd config (which is always risky), a filesystem filling up on the production web server which tanked the search engine for nearly thirty minutes, and a brief outage with Janrain's Engage service, which allows people to use their Facebook ID or Google ID to sign in to the ICPSR web site.

Wednesday, April 11, 2012

March 2012 Web availability

March 2012 was not kind to us.

Clicking the image will display a full-size chart.  But please don't.  It is too ugly.

The main culprit in March was a continuing problem with the reliability of the production web server.  The environment - cooling, electricity, humidity - was fine, and the individual web applications were also fine, but something is not quite right with the kernel.  I think.  (If you would like to join the team as our new Senior Systems Architect and help solve the problem, see my post from last week.)

In March we saw multiple outages, each lasting over an hour.  The script always went something like this:

  1. Load average increases by 5000-10000%
  2. One web application stops responding and logging
  3. KERN.INFO error messages from khugepaged and jsvc appear in syslog
  4. Attempt to restart web application
  5. Fail
  6. Attempt to restart all web apps and their containers
  7. Fail
  8. Attempt to reboot machine
  9. Fail
  10. Optional:  Drive into office if weekend or early morning
  11. Attempt to cycle power on machine
  12. Mix of foul language and prayer
  13. Repeat step #12
  14. Success - machine is working again
Because we use the cloud instead of local, physical servers for many services, and because we haven't had all that many times where the machine needed its power cycled to solve a problem, we don't have things set-up for remote power access.  We'd like to address that.  (If you would like to join the team as our new Senior Systems Architect and help solve the problem, see my post from last week.)

So here's the plan to have a better April:
  1. Disable khugepaged, hoping this might stop the machine from seizing up
  2. Drive faster to the Perry Building, hoping this might result in faster applications of turning the power off and on
  3. Hire the Senior Systems Architect, hoping that having a third pair of eyes on the problem might reveal its true cause and solution
  4. Mix of foul language and prayer, hoping it will ease the pain
And, more seriously, we have also updated a few apps (like Solr) to use local storage rather than NFS-mounted storage for their work, particularly if the app tends to do a lot of writing to the filesystem.  NFS seems to be part of the mystery too.

Monday, April 9, 2012

March deposits at ICPSR

Chart?  Chart.
# of files# of depositsFile format
11230application/msword
93application/octet-stream
8127application/pdf
32application/vnd.ms-excel
42application/vnd.wordperfect
2214application/x-sas
8321application/x-spss
33application/x-stata
22application/x-zip
22image/jpeg
11image/x-3ds
11message/rfc8220117bit
94text/html
52text/plain; charset=unknown
8716text/plain; charset=us-ascii
32text/rtf
11text/x-c; charset=unknown
71text/x-c; charset=us-ascii

A blissfully normal month of deposits.  Usual types.  Usual volumes.

Still need to tweak the automated MIME type detector to stop reporting that it is finding C source code.  The eight files above are most likely plain text files that just happen to have something like a pound-sign or "slash-star" sequence starting in the first column.

Not shown here - because it isn't passing through the deposit system - is a considerable volume of video content from the Gates Foundation.  We have a bit over 6TB that we received in early 2012, and about 1TB of a 20TB collection that will arrive in a steady stream over the next 12-16 months.

If our policy is that the ICPSR deposit system is just one of many mechanisms for ICPSR to accept content, then this seems OK.

But, if we expect the deposit system to be the complete and correct record of ALL incoming content, then we do have a problem.  A 7TB problem that is will grow up to be a big and strong 26TB problem at some point.