Showing posts with label information technology. Show all posts
Showing posts with label information technology. Show all posts

Wednesday, July 17, 2013

ICPSR Web Availability - 2012-2013

Here are the final numbers for ICPSR's web site availability over our last fiscal year:

Click to embiggen
The year did not start off so well, and we reached the nadir quickly.  August 2012 was our worst period of availability in a very bad year for us overall.  January, March, and June 2012 also had very poor numbers.

The main antagonist we faced was a new and unusual problem with our Oracle database server.  For many years we would export the content for backup purposes each evening, and it worked well for a decade.  However, suddenly in 2012 we began to experience an outage just AFTER each export.  Despite intensive analysis by ourselves and local Oracle exports, we never could isolate the root cause of failure.

We eventually "solved" the problem by exporting our database only once per week v. once per day.  That left us more exposed to loss, of course, but it seemed to limit the outages to once per week v. once per day.

We then replaced the hardware with a new machine with a bit more processor and memory, but with blindingly fast solid-state drives. With the new machine deployed we returned to our daily export schedule, and the machine -- and our web availability -- have been in pretty good shape ever since. The machine went into service in April 2012, and the chart above makes it clear that life has been a little less hectic for our on-call engineer since then.

Wednesday, May 1, 2013

EMC anonymous ftp service and transfer_support_materials

I have not seen notes about this in forums and boards, and so thought I would pass this along to others who may be using EMC gear.

About a month ago we had a small problem with one of our NS 120 Celerra NAS units.  (It may have been soft errors on one of its disk drives.)  The Celerra detected the problem, and went to do its usual thing:  collect logs and other analytics, and then copy them to EMC's anonymous ftp site.  Our Celerra uses a utility under /nas/tools called /nas/tools/transfer_support_materials to do this. We noticed that when the Celerra tried to transfer the support materials that too failed.  And this generated an additional series of critical errors.

We logged into the Celerra's control station and ran transfer_support_materials by hand.  And we saw a message like this:

[nasadmin@controller tools]$ /nas/tools/transfer_support_materials -uploadlog
transfer_support_materials[12057]: The transfer script has started.
PING ftp.emc.com (168.159.219.138) 56(84) bytes of data.
From 12.249.233.6 icmp_seq=0 Packet filtered

--- ftp.emc.com ping statistics ---
1 packets transmitted, 0 received, +1 errors, 100% packet loss, time 0ms
, pipe 2
cd: Access failed: 550 Requested action not taken. File unavailable. (/incoming/APM00000000000)
`/nas/var/emcsupport/support_materials_APM00000000000.130407_1351.zip' at 65536 (0%) 49.1K/s eta:5m [Connection idle]

I've replaced our Celerra's serial number with the string "00000000000".

We then ran ftp by hand to see if we could replicate the error:

nasadmin@controller tools]$ ftp ftp.emc.com
Connected to ftp.emc.com.
220-Proceeding further constitutes acknowledgement
to EMC Acceptable Use and Customer Security policies.
Anonymous uploads are immediately moved to a secure server accessible only
within EMC networks.
File downloads from ftp.emc.com are restricted to selected /pub directories, via
temporary secure accounts or via specific permanent secure accounts only.
Anonymous users please login with anonymous and email address as your password
See Powerlink emc278739 for upload instructions.
EMC staff: please refer to current services, FAQ and Best Practices documents at
http://one.emc.com/clearspace/community/active/css/projects/ftp-service
Please email all questions and concerns to ftpquestions@emc.com
220 Please reference the FTP Acceptable Use policy: http://itcentral.corp.emc.com/Policies/AcceptableUse.pdf
534 Command denied.
534 Command denied.
KERBEROS_V4 rejected as an authentication type
Name (ftp.emc.com:nasadmin): anonymous
331 User name okay, need password.
Password:
230 User logged in, proceed.
Remote system type is UNIX.
Using binary mode to transfer files.
ftp> cd /incoming/APM00000000000
550 Requested action not taken. File unavailable.
ftp>

So, the problem was that the directory that holds our support materials (/incoming/APM<serialnum>) was missing or had its mode set to something that disallowed access.

We contacted EMC, and some days later they confirmed that the problem was indeed that the directory was missing, and that they had recreated it.  We then ran ftp by hand to confirm that everything was working again, and it was. That was good news, but when we tried the same thing on our second NS 120 Celerra, we discovered that it too was missing its "support directory" on the ftp server.  So we added that trouble report to our service request, and some days later, EMC confirmed that too had been missing, and then again recreated it.  In speaking with EMC it is a bit unclear if this problem is particular to us or more broad.

The upshot of the story is that if you too run a Celerra or other product that sends support materials to EMC via anonymous ftp, this might be a good day to test out transfer_support_materials to make sure that your "support directory" is intact.  If so, that's great, but if it is missing, you may want to open a service request with EMC soon so that they can recreate the directory for you.  Better to have it in place before your system needs to send support materials, but is not able to do so.

I should note that we're still happy overall with EMC; in fact, we've just purchased the first three nodes of a new Isilon storage system from them.  So the intent here isn't to excoriate them over the missing ftp directory; it was easy to reproduce the problem and to correct it.  But we did wish that we had been able to learn about the problem prior to the disk failure so that it could have been corrected earlier, not when the Celerra was trying to report a disk failure.
 


Friday, September 28, 2012

Be careful when answering, "Both"

When I am working with someone to work through the requirements of a new system or project, I will often ask a series of questions that help shape my understanding of what the person wants.  Often these fall into a pattern where I ask a series of either/or questions, like this:

Do you want it optimized for security or ease of use?

Does this fit into a wholesale or retail delivery paradigm?

Will this be used by external customers or internal staff?

Is this intended to drive new revenues or decrease current costs?

In many ways this is like going to the ophthalmologist who has you look through lens A and then lens B, and then asks the question, "Which was better, A or B?"  Both of us are trying to bring the problem into focus.

The single most dreaded (by me) answer to these questions is: Both.

In some cases this answer really means, "I am not sure what I want."  Or, "I'm too busy to think about this."  Or, "I don't care."

This, obviously, does not help when gathering requirements.  And so it is a real barrier to scoping the project. Sometimes, of course, an answer of Both is a fine start to a longer answer.
We really do need both in this case.  We want to build a system for managing metadata that can be used by both the staff and external people equally easily.  We are changing our entire workflow so that either population can manage our metadata, and this is our new business practice.
That is a fine use of Both.  In fact, if the person gave a different answer, we might needlessly limit the usefulness of the system we build.

And another fine answer, just like we sometimes tell the ophthalmologist is I don't know.  There's nothing wrong with that answer.  However, just like with the ophthalmologist, when I hear this answer, I reach into my bag of lens, and try another pair to bring the issue into focus.

Wednesday, August 1, 2012

ICPSR 2012 technology recap - where did the money go?

A post from a week or so ago showed the sources of money flowing into the technology organization at ICPSR.  This week's post will focus on how that money is spent.

I should note that the focus of this post is on what we call the "Computer Recharge" at ICPSR.  This is a tax paid by all FTEs at ICPSR that is levied on an hourly basis.  Each time an employee completes his or her timesheet and allocates time to a project (exceptions: sick, holiday, and vacation time), a small amount of money also accumulates in the "Recharge."  In FY 2012 this accounted for about 45% of the technology revenue.

Unlike a direct charge for technology where a project or grant may be paying for a dedicated server, extra storage space, or a fraction of a software developer working on custom systems, the Computer Recharge dollars are used to fund technology expenses that benefit the entire organization.  This includes expenses as prosaic as printers and desktop computers, but also includes systems and software development for our delivery systems, ingest systems, and digital preservation systems.

Here's the breakdown in chart form:

No surprise that most of the money goes to pay people:  systems administrators, desktop support specialists, software developers, and a group of very hands-on managers who do a lot of the same work plus project management and business analysis.  This accounts for $862k out of $1218k.

Equipment ($214k) represents desktop machines, new storage capacity, printers, virtual machines, cloud storage, software licenses, and almost every non-salary expense.

Transfers ($103k) is an interesting category.  This is money that we collect as major systems depreciate.  But since the U-M doesn't really "do" depreciation, we use this interesting process instead:


  1. Buy the item (using money from Equipment pot)
  2. Estimate the lifetime of the item (say, five years)
  3. Collect 1/5 of the purchase price each year for five years
  4. At the end of each year, move the 1/5 collected into Transfers
  5. When the item needs to be replaced, use the money in the Transfers pot


So the $103k represents money that was collected in FY 2012 for items that were purchased in earlier years, and which will need to be replaced in FY 2013 or beyond.

All of our other expenses are tiny by comparison:  $19k for Internet access, $7k for telephones and service monitoring, $6k for travel, $5k for maintenance contracts, and $2k for miscellaneous fees and expenses.

One take-away from this is that the essential element in technology budgeting isn't the purchase price of the server, or the annual cost of cloud storage, bur rather the recurring cost of the people who will build, maintain, enhance, and customize the technology portion of the business.

Wednesday, July 18, 2012

ICPSR web maintenance

We're updating a few pieces of core technology on our web server this afternoon:  httpd, mod_perl, Perl, and a few others.  Normally we like to perform maintenance like this during off-hours, but we're doing it at 12:30pm EDT today so that we have "all hands on-deck" to troubleshoot and solve problems.

We've already performed this maintenance on our staging server, and that went smoothly.  Our expectation is that this maintenance will last 15-30 minutes.

Monday, July 16, 2012

ICPSR 2012 technology recap - where did the money come from?

We're putting together some summary numbers for technology spending and investments at ICPSR for FY 2012.  (The ICPSR fiscal year is the same as the University of Michigan's, and runs from July 1 to June 30.  We've just recently closed FY 2012.)

The first set of numbers shows the allocation of effort in FY 2012 by funding source. The unit of measurement in this pie chart is HOURS (not DOLLARS) that were expended in FY 2012 by each funding source.  (We originally wanted to calculate dollars, but that turns out to be an even bigger effort.)  Here's an interactive chart:



This is an interactive Google Docs chart.  If you click slices of the pie, it will identify the funding source.

The main source of technology effort funding comes from the Computer Recharge, an hourly "tax" that ICPSR levies against all projects.  Although it is one single funding source (nearly 45% of hours worked in FY 2012 were billed against this source), I have split it into two sub-categories, one for what I am calling "IT" and one for "SW" (software).

The "SW" portion includes the effort of all staff who are professional software developers.  The type of work performed by this team using this account includes enhancements and maintenance for ICPSR's core data curation and data management systems, and investments in new products and services such as software developed to support our IDARS system for applying for access to datasets.

The "IT" portion includes the effort of the remainder of the staff which tends to include systems administrators, architects, network managers, and desktop support specialists.  I also allocate my own time to this bucket since the majority of my non-contract, non-grant effort over the past year has been in building and architecting technology systems.

Other big slices of the "IT pie" include the work of staff members who are explicitly funded by projects such as our CCEERC and RCMD web portals; our two Bill and Melinda Gates Foundation grants; the ICPSR Summer Program, and many more.  In fact, there are over 20 separate funding sources used to support technology at ICPSR; the pie chart shows 18 because I grouped several small ones into a category called "Misc."

If this gives the impression that there are many, many projects and activities at ICPSR that involve technology, that's good!  That is certainly the case.

However, if "focus wins" then we're in a little bit of trouble.  My sense is that each of this 20-some funding sources has at least one unique project with its own business analysis and project management needs, and it is sometimes the case that different projects have antithetical technology needs.  I see this play out in all phases of the OAIS lifecycle.  ("I want you to build a system that makes it as easy as possible to fetch datasets from ICPSR" v.  "I want you to build a system that requires significant effort and oversight to fetch datasets from ICPSR.")


Friday, June 8, 2012

May 2012 web availability

Web availability was good, but not great, in May 2012:

Click to enlarge


Five main episodes of 25 to 59 minutes account for almost all of the 242 minutes of unavailability in May 2012.

On May 13 the production web server seemingly lost power, and it required an additional reboot and some TLC to bring the system back on-line fully (33 minutes).

On May 18 the search index on our CCEERC web portal became corrupted, and that disabled much of the usefulness of the site for nearly an hour (49 minutes).

On May 22 we saw the first of two episodes where the proxy (AJP) between Apache httpd and Apache tomcat faulted.  This did not recover on its own and required some help from the technology team.  This resulted in a medium-duration outage (27 minutes), and another similar fault occurred on the evening of May 31 (24 minutes).

On May 31 our production database server faulted, requiring a manual power-cycle, and also requiring the production web server to be rebooted (59 minutes).

My sense is that while we're in much better shape with regard to the instability caused by khugepaged, we are starting to see something a little amiss with the Apache proxy system.  It isn't clear to us at this time if the issue is faulty software or sub-optimized configuration on our part.

Monday, May 28, 2012

Seven tips to survive Going Google

Are you Going Google on your campus?

Michigan is going Google in 2012.  A few of us migrated from existing IT systems at the University of Michigan in January, and so have been living the Google dream, but needing to still work with our colleagues who are still using legacy systems for email, documents, calendars, etc.  I've seen some of the complaints people have had as they have moved to Google, and I can also see some of the future problems ahead.  You can save yourself and your colleagues hours of frustration by following a few simple rules.  And so I present the tech@icpsr survival guide to Going Google.

One, stop organizing your email.  You don't need to spend your time that way any longer.  The only reason you  needed to do that in the old world was that you had a low quota for storage (and so kept moving folders off of the mail server and on to local storage) and you have a bad search.  Now you have plenty of storage and a great search.

Two, stop asking people when they are available via email.  Look at the shared calendar.  If you cannot see the person's calendar, tell them to fix the access controls.  And if they won't, then make them schedule the meeting instead.

Three, never, ever download a Google Doc and start editing it in Microsoft Office.  Once you move the document out of Google Docs and into Office you break sharing, introduce odd formatting, make it difficult or impossible to fold the changes back into Google Docs, and commit other crimes against documents.

Four, do not send documents as attachments.  Make a Google Doc.  Share it with your collaborators or readers.  Do not fill up their Gmail allocation with your documents.

Five, use Chat for the quick stuff.  Got a quick question?  Need a real-time response?  Stop using email. Got a long question?  Do not need a real-time response. Use email.  Long chats are just as bad as 4-minute voice mails.

Six, stop doing THAT in email.  If you find yourself encountering barrier after barrier trying to execute a business process via Gmail, it is likely that email is simply the wrong solution.  Need a shared archive of email?  Use a Google Group.  Need a help desk, ticketing, or request system?  Use Footprints or JIRA or any one of many open source or hosted solutions.  Need a place to share and edit a catalog of information?  Use a Google Site.  Many of the Gmail-related headaches I've seen on campus are caused when people are trying to use email as a substitute for a more complex business process.

Seven, get a personnel email address NOW.  My experience is that it is always risky to rely upon an employer or a telecom to supply your email.  People who were using their @umich.edu email address for personnel use and using their @department.umich.edu for work are now in a pinch at UMich.  The @umich.edu address is necessarily becoming the one for work use, and they are now scrambling.  Don't wait, go get a personnel Gmail or Yahoo or Hotmail or other email address and mail account today.

Wednesday, May 23, 2012

How many bits of video will I stream?

We have a copy of video preservation and access projects for the Bill and Melinda Gates Foundation.

One project consists of highly restricted video content, and we believe the demand will be low enough - dozens or fewer of simultaneous video consumers - that we can stream the content quite comfortably from ICPSR.  (ICPSR shares a 1 Gb/s network pipe with one of the other centers at ISR, and the bit-rate of each video is about 700 Kb/s.) A follow-on project consists of less restricted video content that we believe will have broad appeal.  A key question for the IT director is if the demand will be so high that it will exceed our capacity to deliver.

My colleagues are projecting that we will have peak simultaneous usage of 2000 video consumers.  A little back of the envelope math (total consumers x 700Kb/s) makes it clear that our network pipe is too small; we'll need to move the content elsewhere for delivery, or split the load across several network locations to make delivery feasible.  Unfortunately this collection is quite large - 20 TB - and so making lots of copies to spread the delivery across lots of locations will be expensive.

Another approach is to move the content into a content delivery network (CDN).  In this scenario the CDN operator will charge us a fixed rate to store our content and a variable rate to stream our content.  So how much will all this cost?

The storage is easy.  We have 20 TB, and so we can calculate the storage costs quite easily.The streaming costs are more tricky, however.  Typically one's costs are tied to the total number of bits streamed each month, but our only data point is the maximum number of total simultaneous video consumers.  So how do we calculate the expected cost?

We've been struggling with this for a while, and I don't know that we've hit upon a good solution.  But we do have A solution.  Here it is....

What if we were to graph the number of concurrent video consumers?  And what if we assume that the graph will be a curve, a Gaussian curve in particular?

Source: NIST
Our Y-axis can measure the total number of simultaneous video consumers at a given point in time.  We have our maximum height value (2000) as one data point.

Our X-axis can measure time of day where each point is a single second in a 24-hour period.  And we'll choose the starting point and ending point so that the maximum height falls in the exact middle of the graph.

If we calculate the area under the curve this will tell us the total number of consumer-seconds, and we can then multiply that by 700 Kb/s to calculate the total number of kilobits streamed in a 24-hour period.  And we can divide by 8 x 1024 x 1024 if we want to turn kilobits into gigabytes, a standard unit of measurement for calculating streaming costs.

To calculate the area under the curve we need to know the maximum height (2000) and we need to estimate how "fat" or "thin" our curve will be.  (This is related to standard deviation in a normal distribution.)  So if our X-axis is seconds, we might pick something like 60 (for a very pointed curve) or 3600 (for a flatter curve). And if we call the height 'a' and the width 'c' our formula for measuring the area (bits) is:

a x c x SQRT ( 2 x PI )

We can then use fixed rates to turn number of consumers into number of GBs.  And if we have a per-GB price, we can turn that into a total daily cost.  I made a little calculator at Zoho to help with this.  (Note that you must be sure to use the Tab key to move through the form.  Hitting the Enter key or using the Submit button stores the information in a throw-away table at Zoho and clears the form. )

For example, if I think I'll have a maximum of 2000 simultaneous consumers (a = 2000), and I think my curve will be medium width (c = 1800), and my video is 700 Kb/s, and my price to stream is $0.25/GB, then my daily cost will be approx $188.

Monday, May 7, 2012

OpenSSL FIPS 140-2 and RHEL 5

Since I often use the Internet to find guidance and answers to questions, I thought I'd add my own small contribution back to the community.

We found ourselves needing to build a FIPS 140-2 compliant version of the OpenSSL openssl command-line utility recently.  A good starting point is the OpenSSL FIPS 140-2 User Guide.  It contains instructions for where to find OpenSSL source, and importantly, instructions for verifying the integrity of the distribution, and this is a necessary component of building a FIPS 140-2 openssl.

I worked through the guide through nearly page 23.  However, when I reached section 4.2.1, things started to go wrong for me.  I was able to run config with no problem, but the make failed with an error about not having a target for fipscannister.o.

I then found a very helpful bit of advice in a Google Group post about OpenSSL and FIPS 140-2.  But that advice didn't quite match our environment (it was Ubuntu, and we run RHEL).  So here are my directions for building an openssl utility on RHEL 5.

First, in our case we downloaded and verified the integrity of openss-1.2.3 using the directions from the guide above.  Here are the results in a little sandbox:

batch-bryan:; pwd
/tmp/openssl-fips
batch-bryan:; ls -R
.:
lib/  src/

./lib:

./src:
openssl-fips-1.2.3.tar

Now, we head into the src directory, and unpackage the tarball:

batch-bryan:; cd src
/tmp/openssl-fips/src
batch-bryan:; tar xf openssl-fips-1.2.3.tar

And then configure things to build the FIPS canister and utility:

batch-bryan:; cd openssl-fips-1.2.3
/tmp/openssl-fips/src/openssl-fips-1.2.3
batch-bryan:; ./config fipscanisterbuild --prefix="/tmp/openssl-fips"

Now to build and install:

batch-bryan:; make
batch-bryan:; make install

And there it is:

batch-bryan:; ls /tmp/openssl-fips/lib
engines/             fips_premain.c       libcrypto.so@        libssl.so@
fipscanister.o       fips_premain.c.sha1  libcrypto.so.0.9.8*  libssl.so.0.9.8*
fipscanister.o.sha1  libcrypto.a          libssl.a             pkgconfig/


And there is the utility:

batch-bryan:; cd /tmp/openssl-fips/bin
/tmp/openssl-fips/bin
batch-bryan:; ls
c_rehash*  openssl*
batch-bryan:; setenv OPENSSL_FIPS 1
batch-bryan:; ./openssl version
OpenSSL FIPS Object Module v1.2

But if we try the stock one:

batch-bryan:; /usr/bin/openssl version
13789:error:2D06C06E:FIPS routines:FIPS_mode_set:fingerprint does not match:fips.c:493:
batch-bryan:; unsetenv OPENSSL_FIPS
batch-bryan:; /usr/bin/openssl version
OpenSSL 0.9.8e-fips-rhel5 01 Jul 2008

Not so much.

To build and install the FIPS 140-2 compliant version of OpenSSL in a more "real" location than a sandbox, just change the value used in the prefix variable used in the config invocation above.





Friday, April 20, 2012

It's official - we have no love for khugepaged

As web site visitors - and the IT staff - experienced through February and March after we upgraded to RHEL 6, life was tough.  Very tough.



We saw two months of ghastly web service availability, well below the 99% goal.  Lots of pages in the middle of the night.  Lots of trips to the office at all hours and all days to cycle power on the server.

It was clear that khugepaged was involved somehow.  Was it the victim of something else?  Or the cause?

Based on the most scanty of evidence and great desperation we disabled khugepaged on March 29.  And since then?

[ sound of knocking on wood ]

The machine is back to its old self.  One very short-lived (seven minutes) outage based on a bad rewrite rule that we added in response to a request, and then had to back-out.

Who knew that this simple command:

root# echo never> /sys/kernel/mm/redhat_transparent_hugepage/enabled

could generate so much happiness?

Wednesday, April 18, 2012

Great FLAMEing file identification service

Some parts of the FLAME project will lend themselves to a microservices approach.  Microservices, like cloud computing, is a trendy, useful concept, but without a crystal clear definition.  But my take is that a microservice is something that performs one small, but useful bit of work, and which can be swapped in and out of an overall architecture at a component level.  It needs to have very clear inputs and outputs, and cannot contain any "secret sauce" that isn't part of its functional role.

Do not try this street magic at home.
One common activity at ICPSR is automated file identification.  Historically we've done this with the venerable UNIX utility file, but where we modify the magic database heavily, particularly for the formats we see most often.  We also post-process the output from file where we need additional handling above and beyond the capabilities of the magic database (e.g., making decisions based on the name or extension of the file).

Managing the magic database is not for the faint of heart.  (Try updating the Vorbis section.)  And this management has gotten both harder -- RHEL 6 uses a new format for its magic database which is incompatible with RHEL 5 -- and easier -- the new format eliminates the pesky magic.mime database.  However, we've gotten reasonably competent at managing magic and have come to rely on it for file format identification.

In support of the FLAME project we even created a little web service that takes a file's content and its name as input, and delivers a little snippet of XML as the output.  The XML contains the "human readable" answer from our magic database and the "MIME type" too.  This is our first FLAME-inspired web service.

If you'd like to try it, you can use your favorite form-capable URL transfer utility to do so.  Here's an example where I have run curl on one of our RHEL machines:


dhcp-bryan:; curl -F "file=@uuid-comparison.xlsx;filename=uuid-comparison.xlsx" www.icpsr.umich.edu/cgi-bin/wsifile<?xml version="1.0" encoding="utf-8"?><wsifile><ifile>Microsoft Excel</ifile><ifilemime>application/zip; charset=binary</ifilemime><uploadInfo>application/octet-stream</uploadInfo></wsifile>

feeding in an Excel file as the input, and another with a plain text file:

dhcp-bryan:; curl -F "file=@/etc/resolv.conf;filename=resolv.conf" www.icpsr.umich.edu/cgi-bin/wsifile<?xml version="1.0" encoding="utf-8"?><wsifile><ifile>ASCII text</ifile><ifilemime>text/plain; charset=us-ascii</ifilemime><uploadInfo>application/octet-stream</uploadInfo></wsifile>

and an interesting MS Word file:

dhcp-bryan:; curl -F "file=@2011-03CouncilPandAminutes.doc;filename=2011-03CouncilPandAminutes.doc" www.icpsr.umich.edu/cgi-bin/wsifile<?xml version="1.0" encoding="utf-8"?><wsifile><ifile>CDF V2 Document, Little Endian, Os: Windows, Version 5.1, Code page: 1200, Number of Characters: 0, Name of Creating Application: Aspose.Words for Java 4.0.3.0, Number of Pages: 1, Revision Number: 1, Security: 0, Template: Normal.dot, Number of Words: 0</ifile><ifilemime>application/msword; charset=binary</ifilemime><uploadInfo>application/octet-stream</uploadInfo></wsifile>

Feel free to try it out, and to post reactions, suggestions here.

Monday, April 2, 2012

FLAME update

We have three different tracks running on FLAME.

One track is conducting an analysis of the business requirements ICPSR has for what we are calling our "self archived" collection.  This is a collection of material best represented today by our Publication Related Archive, a set of materials that receives very, very little scrutiny between time of deposit and time of release on the web site.  We are imagining a future world where the quantity of "self-archived" materials increases dramatically from today's volumes, driven by NIH and NSF requirements to share and manage data.

I see the following questions generating the most discussion on this track:  How much disclosure review is necessary before releasing the content publicly?  Should the depositor have "edit" access to the metadata?  If so, should it be moderated or completely open?  How much "touch" does ICPSR need to have on these materials?

Another track is working on a crisp, concrete definition of what it means to "normalize" a system file from SAS, SPSS, or Stata.  ICPSR has long said that our approach is to "normalize" such files, producing plain ASCII data and set-ups, but what does that really mean?  And is that really possible?

I see the following questions generating the most discussion on this track:  Is ASCII the right thing, or ought it be a Unicode character set?  Are set-ups the right documentation or should it be DDI XML?  If we choose the former, is it sufficient to produce set-ups compatible with the original content type (e.g., SAS setups for a SAS file)?  What about precision?  Length of variable names?  Question text?  Is it possible to normalize without loss, and if not, how much loss is acceptable?  Can a computer do this without human intervention 99% of the time?

And the last track is working on a matrix that maps a set of parameters (inputs) to a resulting preservation commitment and set of actions.  For example, if one has a file which contains "documentation" (the type of content) in XML format in the UTF-8 character set (the format of the file), then perhaps the preservation commitment is "full preservation."

The key questions here, I believe, will be around what the right list of parameters is.  And if any of the parameters uses a controlled vocabulary, what's in the CV?  And what exactly does it mean to have a "full preservation" commitment?  What's involved beyond just keeping the bits around, which is presumably all one does with "bit-level preservation?"

Wednesday, March 28, 2012

The cloud is not a hard drive

I recently read a piece in one of the IT trade publications about why buying infrastructure via a cloud provider (Amazon in particular) was a bad deal, and how one would be much better off buying storage in-house.  As is often the case, the opinion-piece compares apples to oranges, and the analysis is so flawed, one can't really make any use of the conclusion.  However, my experience is that the piece is hardly unique in its flaws, and that motivated today's post.

The usual analysis is to compare the cost of storage when purchased as a consumer-grade disk drive v. storage when purchased via a cloud-based infrastructure provider.  The numbers -- $100 for a 2TB drive from Best Buy v. $100/mo for 1TB of S3 space at Amazon -- are thrown out as directly comparable, and thus it is plain to see that it is better to spend that $100 just once to get storage rather than spending the same amount of money each month (and for less storage to boot).

The analysis then usually performs some hand-waving about how disk drives sometimes fail (really?) and how they don't last forever (shocking!), but these facts are just in the noise, and it really is clear how much better one is purchasing storage rather than "renting" it in the cloud.

However, this sort of comparison is like looking at the cost of one's grocery bill for a given meal v. the cost of having a meal at a restaurant.  The first is merely one part of the second, and they simply are not the same thing.

A disk drive from Best Buy can be a very useful thing.  I myself own one of these very same drives, and I use it every couple of weeks to back-up a PC we have at home.  (Clearly my disaster recovery process is not good.)  Because the storage only needs to be "lit" every so often, and only needs to be accessed from a single location, it works well.

However....

If instead I needed my storage to be available 24 x 7, and to have some credible DR plan, this wouldn't do.  Or if I wanted to be able to use a lot more storage some days (or weeks or months) and less storage on others, then buying storage for peak isn't so good.  Also, when it is "lit" my storage consumes some electricity to operate.  So maybe an analysis that includes things like mirrored copies, electricity, front-end access systems, the network, and .....  gets to a more apples-to-apples comparison.

The cloud may be the best solution to some problems, and a dreadful solution to other problems.  Maybe way overpriced.  But a lot of different stuff makes up "cloud storage" and the underlying media is only one of many; in fact, it may be the least expensive component.

Wednesday, March 21, 2012

Zynga says "buh bye" to Amazon

So one big bit of news is how Zynga, the company responsible for such Facebook games as Farmville and Words With Friends, is building out its own IT infrastructure (zCloud) to host its games rather than continuing to rely solely upon Amazon for this infrastructure.

The obvious question:

Why?


Building your own cloud is a big bet:  data centers, racks and racks of equipment, servers, network switches, cables.... lots of cables.  That's a very big investment to make for a platform that's just overhead to your main business: developing software (games).

However.....

If the platform is actually your business, then it makes a lot of sense.  In that world the product isn't software; the product is the platform.  In that world you want to own the platform.  You want the control (costs, performance, etc).

Zynga says that it will still use AWS for spikes in service that it cannot service within its own zCloud platform, and so it isn't the complete end of the relationship for the two companies.  But it is still a very big change for Zynga.

Wednesday, February 29, 2012

Sixteen products or one?

A recent conversation with Nathan Adams, ICPSR's Assistant IT Director for Software Development got me thinking about this....

It's no secret that ICPSR uses a package called Survey Documentation and Analysis (SDA) from UC Berkeley as our on-line analysis system.  But people may be surprised to learn that this one product forms the underpinnings of more than a dozen closely related ICPSR on-line analysis products.

One, Anonymous Analysis : This is where we make a dataset available via SDA and there is no authentication allowed.

Two, Authenticated Analysis : One must authenticate using MyData, Google, or Facebook.

Three, Member Analysis :  One must authenticate and also be using a computer located on the campus (even virtually) of a member institution.

Four, Private Analysis : One must authenticate and the identity used must be a member of a previously created group of identities.

Five through eight, Secure Analysis : Like any of the options above, but where the raw, proprietary, binary data files reside on a separate server, and where the ICPSR web server accesses the content via HTTPS rather than through the filesystem.

Nine through Sixteen, Non-disclosed Analysis : Like any of the eight options above, but where SDA's disclosure.txt controls have been used to attempt to prevent unintentional disclosure.

So sixteen different combinations!  And it is easy to imagine even more cropping up in the months ahead.

My experience is that one ends up with sixteen different online analysis "products" when things grow organically over time.  When things evolve due to a small tweaks in response to requests like, "Hey, could we use SDA for this, but with just one small change ..... ?"

It is easy to see how it happens.  But when things grow over time like this, they end up suffering from a profound lack of design, and end up costing more to maintain.  They are fragile.  They break when you change things, like the hardware.  Or the OS.  Or the NAS.  Or the authentication scheme.  Or the oil in your car.

So probably time to pull back a bit, pull together a team of content owners, and start asking some questions.

If we were going to start fresh today with an on-line analysis system, what should we build?

What sort of access controls are needed to prevent bad guys from using it?

What sort of disclosure mitigation capabilities are required to prevent accidents from happening?

To which populations might we need to restrict access?

What does the user experience look like?  Is this geared for the novice or for expert-in-a-hurry?  Or do we have multiple audiences and so need to build more than one experience?

Time to design.

Wednesday, February 22, 2012

Going Google

The University of Michigan is rolling out Google Apps for Education throughout 2012.  A few of us are in the early (1.0) pilot population, and this group made the jump from a variety of legacy University of Michigan email and calendar systems on January 16, 2012.  I first reported on this new initiative late last year, and it's now time for an update.

I should note that I have been using Google's productivity tools outside my professional life for many years, and so there is not much of a learning curve.  I think this will also be true of some of the more broad population at UMich, but will not be true universally.  And I should also note that I had been using a second Gmail account for my professional life too for the past 2-3 years.  The main driver for me was storage space.  While I'm not a huge fan of Outlook and Exchange, the service operated by ICPSR's parent organization - the Institute for Social Research - was always solid.  However, the killer was that the allowable quota for mail was very low (400MB by default), and so I found it frustrating to always be shuffling email off into either the Trash Can or into PST mailboxes.  It was especially rough when it came time to search for something.

And so the move from a consumer Gmail account that I use for work to a Google Apps Gmail account that I use for work has been a small change.  The change from Exchange to Google Calendar for managing meetings has been a bigger change.  On the plus side I'm finding it much easier to manage a single, coherent picture for meeting invitations; I had been trying to manage everything inside of Exchange before.  However, I probably receive 100 meeting invitations for every one I generate myself, and so I haven't had to spend much time and effort ensuring that meetings I create on my Google Calendar are ending up on the ISR Exchange server intact. In fact, most of the headaches I experience with calendaring are related to cases where someone generates an invite within the ISR Exchange server, but does not include anything in the "body" of the invite.  If I try to "read" the invite on a mobile device (e.g., Safari on an iPad), the meeting invite shows up as an empty message.  And so I then track down a "real" computer to see what the meeting invite is all about.

My main take-away so far is that moving from Exchange to Google would best be done (1) quickly, and (2) all at once.  My sense is that we early adopters will continue to face a few headaches like above until the rest of the organization moves to Google in 3-6 months.

Wednesday, February 8, 2012

Network maintenance - Sunday morning (EST) Feb 12, 2012

The University of Michigan central IT organization, ITS, will be upgrading the network gear that connects ICPSR's building to the campus data network.  The work is scheduled to start at 6am (EST) on the morning of Feb 12, 2012 and should take between one and two hours.

The ICPSR IT team will redirect traffic from the production system to our cloud replica during the maintenance period.  The replica runs in Amazon's cloud and features services such as search, analyze, and download, but purposefully does not enable features such as deposit.

Between the cut-over to the replica, the network maintenance, and the fallback to the production system, access will likely be a little rocky next Sunday morning.  Like a freeway during construction, if it is possible to take a detour around ICPSR's web site on Sunday morning, that's the safest route.  But if you find that you need to download some data or use the site, the replica will be available.

Monday, February 6, 2012

Job posting - again

We're posting a job description for a senior software developer for a third time.  If there is a recession in the IT business in SE Michigan, somebody forgot to tell our pool of potential applicants.

This position is very much like other software developer positions at ICPSR.

In practice there is a blend of business analysis, system design, software development, and second-line on-going support for the stuff you write.  Building stuff in java to run under tomcat is a must.  Experience with Oracle, Eclipse, and one or more frameworks is very useful.

The main support for this position is our Bill and Melinda Gates Foundation MET Extension project.  This is a two year grant to build systems that will delivery video content to researchers in the social sciences and education.  We've had a lot of success - like a 0% failure rate - at hiring people to work on project X, and then moving them over to project Y a few years down the road.  Project X is in this case is MET Extension; project Y is unknown.  But so was the MET Extension project 10 months ago...

If you have interest in the position and would like more info, please feel free to drop me a note.  And here's a short-lived, but direct, link to the job http://umjobs.org/job_detail/66372/software_developer_senior

Monday, January 30, 2012

Customer service, Zingerman's style

Our parent organization, the University of Michigan's Institute for Social Research (ISR), is working with the training component of the Zingerman's family of companies - ZingTrain - to build a customized training module for use at the ISR.  The focus is, of course, on delivering excellent customer service, and I had the opportunity to attend a session led by two ZingTrain consultants.

I don't want to give away too much of their "secret sauce" but I found their interaction with the group engaging and informative.  I almost used the word "presentation" but that feels wrong; it really isn't a monologue whatsoever.  And there are no Powerpoint slides in sight.  As you might expect the ZingTrain folks shared some tips and techniques about how they build the right culture and right processes.  And they brought goodies from the Bakehouse!

I started to think about some of the tips and techniques I've learned to use in the technology business over the years.  In this realm an awful lot of the interaction with others takes place electronically, and so one doesn't have all of the visual cues and tonal cues one normally can use in conversation.  For example, how do you let someone know that if the solution you have offered does not work, you want and expect the person to let you know so that you can keep trying to solve the problem?  How do you let them know that you will own the problem until it is solved?

One easy way, of course, is to be explicit.
If that doesn't do the trick, please let me know.  I have a few other ideas we can try.
By asking the person to return and letting them know that "we" can try some other things, it shows that one is engaged.  It lets them know that this is the start of a conversation, not the end of one.

On the other hand, I will often see people write this instead:
Hope this helps.
I know people often write this with the best of intentions, but consider how people may read it.  It sounds like the conversation is over.  "Here, try this.  I hope it works.  But if it doesn't, it's your problem, not mine." There's no invitation to come back for more advice, more assistance, more analysis if the issue hasn't been resolved.

And that's my customer service tip for the month.

Hope it helps. :-)