Showing posts with label video. Show all posts
Showing posts with label video. Show all posts

Monday, April 29, 2013

ICPSR launches Measures of Effective Teaching web site

Some of my colleagues, including ICPSR Director George Alter, gave a demo of one of our newest Web sites and collections at the American Educational Research Association 2013 annual meeting on Sunday.
MET LDB web site
Click the image to navigate to the live site
My team has built the video portal portion of the system.  The portal enables a researcher to play a list of videos that s/he has selected to view based on an analysis of the associated quantitative data and tagging data.  Access to the video and datasets is restricted and requires one to complete a data use agreement via ICPSR's web-based request system.

We're grateful for the support we've received from the Bill and Melinda Gates Foundation to make all of this possible.

Wednesday, October 3, 2012

Setting up Kaltura - part VI

We'll focus on the Kaltura Drop Folder feature today.  The Drop Folder offers a mechanism whereby an enterprise can bulk upload content without human intervention.  In principle this is an excellent way for a library or archive to ingest many objects into Kaltura without some poor archivist performing individual (or group) uploads via a web GUI.  In practice the mechanism works smoothly when things are going well, but it can be a little difficult to diagnose problems when things go awry.

For example, here's a sample display from the Drop Folders panel from our Kaltura Management Console (KMC), which serves as an all-in-one dashboard for managing content:

Click to see a full-size image

According to this display we have just ingested three items:  an XML file and two video files.  In this particular case the XML file contained all of the metadata for the two video files, and contained instructions that told Kaltura that these were new items to ingest.  The Status field shows a value of Done, and the Error Description field is empty.  This seems good.

We can also see status information if we navigate to the Upload Control panel and select the Bulk Upload view.  Here we see similar info:

Click to see a full-size image
Again, this seems like good news.  The Notification column shows a value of Finished successfully.  Hooray!

But not so fast, my friend....

If we examine one of the video files under the Content panel (Entries tab), we see that none of the extended metadata is present.  We can see the Custom Data fields, but they are all empty.  Hmm, what happened?

If we navigate back to the Upload Control display, the last column offers some possible help:


There is an Action available to download a log file.  That sounds promising.  Let's do that.

The log file is in XML format, and if we open it up in a good browser or text editor or XML editor, we find XML that looks very much like the ingest XML we used in the Drop Folder.  And if we scroll all the way down to the bottom, we find this snippet:

<item><result><errorDescription>customDataItems failed: invalid metadata data: Element 'METXVideoSubmissionElectronicBoardUsed': [facet 'enumeration'] The value ' ' is not an element of the set {'Y', 'N'}. at line 87 Element 'METXVideoSubmissionElectronicBoardUsed': ' ' is not a valid value of the local atomic type. at line 87 </errorDescription></result>
This is telling us that we messed up one of the metadata fields.  If we look at the original ingest XML and find the statement that is supposed to be setting METXVideoSubmissionElectronicBoardUsed, sure enough, there is no value.  (The error occurs on line 296, not 87, which is a bit confusing.)

So the good news is that if we notice the error, we can find a log that will point us at the error.  But detecting the error is a little tricky, and it is easy to see how this would be difficult if we were ingesting, say, 100 items at a time.  So this is not awful, but is also not quite as nice as we might like.

Suggestions:
  1. If the XML contains Custom Data, and the Custom Data has errors, but the video still ingests, perhaps a Status of something like "Done with errors" (in the Drop Folders display) or "Finished with Custom Data errors" (in the Bulk Upload Log display).
  2. Make the diagnostic message (errorDescription) available without needing to download a file.  This could appear in a new column, or perhaps in a text pop-up.
  3. If N - 1 elements of the Custom Data are good, but one is bad, it would be nice if the other N - 1 Custom Data fields are set.  That would make it possible to correct the error manually in the KMC rather than copying fresh XML into the Drop Folder.
  4. Suppress the line numbers since they are relative to the log file XML, not the original XML.
Again, overall the Drop Folder feature is very nice, and we will indeed use it to ingest the 20-30,000 video files in our collection.  But since it is likely that we will sometimes make a mistake within the XML (say, forgetting to escape a certain character), it would be great if the KMC would make it hard to detect and diagnose mistakes.

Wednesday, September 26, 2012

Kaltura pilot at the U-M and ICPSR

The University of Michigan in-house newspaper/newsletter ran a nice piece on the Video Contement Management pilot (using Kaltura), which includes our Bill and Melinda Gates Foundation project:  Measures of Effective Teaching - Extenstion.

Here is the part about us:
• ISR and the School of Education are engaged in a collaborative research project in which a large collection of video assets is a primary data-type. This data will become a shared repository available to research partners at universities, public agencies, and private foundations.
The timing here was perfect for us since we were in the market for a system to manage and stream about 20TB of video to thousands of simultaneous users.

Friday, September 21, 2012

Setting up Kaltura - part V

In this post we will look at the XML we use to ingest content into Kaltura through its Drop Folder mechanism.  To re-cap an earlier post about the Drop Folder, ours is a subdirectory under the home directory of the 'kaltura' user.  We provisioned this account on a special-purpose machine that Kaltura accesses via sftp to fetch content without human intervention.

Kaltura has a pretty nice guide to building the XML, which looks an awful lot like Media RSS.  And they also make some short examples available, but we always find it useful to have a real-world example.  Here's ours.

Preface: Our stuff is a little unusual.  That is:
  1. We always have pairs of videos, one classroom and one blackboard
  2. We have lots of metadata and it applies to both videos, and so the metadata gets repeated in the XML.  I have excised much of the metadata in our XML for this post
  3. All of the metadata is fake; it is not real metadata about an actual classroom video
  4. You can find a copy of the XML that we diagram below at this URL http://goo.gl/OxbyU
  5. Our use of Kaltura is in support of the Measures of Effective Teaching (Extension) project, and so there are many references to 'metext' in the metadata
  6. We will be generating the XML for ingest programatically
Here goes.

 <?xml version="1.0"?>  
 <mrss xmlns:xsd="http://www.w3.org/2001/XMLSchema" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="ingestion.xsd">  
     <channel>  
         <item>  
             <!-- These do not change -->  
             <action>add</action>  
             <type>1</type>  
             <!-- This changes for each video -->  
             <referenceId>12345-Board-video-rendition-MP4-H264v-Standard.mp4</referenceId>  
             <!-- This does not change -->  
             <userId>metext</userId>  
             <!-- This changes for each video -->  
             <name>RCN 12345 board video</name>  
             <!-- This changes for each video -->  
             <description>RCN 12345 board video</description>  

We have the usual XML stuff at the beginning, and then the start of the Media RSS.

Action = add for new content.

Type = 1 for video content.

ReferenceID is the name of the original file.

UserID is the pseudo-user who will be linked to the content in Kaltura.

Name and Description are exposed as base metadata in Kaltura.

             <!-- Always assign two tags, one called metext and the other board or classroom -->  
             <tags>  
                 <tag>metext</tag>  
                 <tag>board</tag>  
             </tags>  
             <!-- This does not change -->  
             <categories>  
                 <category>metext</category>  
             </categories>  
             <!-- This does not change -->  
             <media>  
                 <mediaType>1</mediaType>  
             </media>  

We tag everything with the project name and indicate whether it is a video of the blackboard or the classroom.

Kaltura uses Categories as a main way to browse and find content.  We treat this as if it were a type of "is in collection" sort of attribute.


MediaType = 1 for video.

             <!-- This changes for each video -->  
             <contentAssets>  
                 <content>  
                     <dropFolderFileContentResource filePath="12345-Board-video-rendition-MP4-H264v-Standard.mov"/>  
                 </content>  
             </contentAssets>  

This tells Kaltura that the video file is in the Drop Folder along with the XML.

Now for our project-specific metadata, which fits into a Kaltura structure called Custom Data:

             <!-- This changes for each video -->  
             <customDataItems>  
                 <customData metadataProfileId="22971">  
                     <xmlData>  
                         <metadata>  
                             <METXDistrictDistrictName>Ann Arbor</METXDistrictDistrictName>  
                             <METXDistrictDistrictNum>20</METXDistrictDistrictNum>  
                             <METXSchoolSchoolName>Huron High School</METXSchoolSchoolName>  
                             <METXSchoolSchoolMETXID>33</METXSchoolSchoolMETXID>  
                             <!-- Kaltura wants the date to be in xs:long format, that is, seconds from the epoch -->  
                             <METXVideoSubmissionCaptureDate>1344440267</METXVideoSubmissionCaptureDate>  
                         </metadata>  
                     </xmlData>  
                 </customData>  
             </customDataItems>  

The ID attribute is from our Kaltura KMC.  I had to create the Custom Data schema first, and then reference it in the ingest XML here.

Most of the metadata fields are simple strings or strings from a controlled vocabulary.  We do have one date item, and sadly Kaltura expects it to be in a difficult-to-use format, seconds since the epoch.

After this section of the XML is a closing tag for item, and then the whole thing repeats with only minor variation for the classroom video.  I'll include it below for completeness.


         </item>  
         <item>  
             <!-- These do not change -->  
             <action>add</action>  
             <type>1</type>  
             <!-- This changes for each video -->  
             <referenceId>12345-Classroom-video-rendition-MP4-H264v-Standard.mp4</referenceId>  
             <!-- This does not change -->  
             <userId>metext</userId>  
             <!-- This changes for each video -->  
             <name>RCN 12345 classroom video</name>  
             <!-- This changes for each video -->  
             <description>RCN 12345 classroom video</description>  
             <!-- This changes for each video -->  
             <tags>  
                 <tag>metext</tag>  
                 <tag>classroom</tag>  
             </tags>  
             <!-- This does not change -->  
             <categories>  
                 <category>metext</category>  
             </categories>   
             <!-- This does not change -->  
             <media>  
                 <mediaType>1</mediaType>  
             </media>  
             <!-- This changes for each video -->  
             <contentAssets>  
                 <content>  
                     <dropFolderFileContentResource filePath="12345-Classroom-video-rendition-MP4-H264v-Standard.mov"/>  
                 </content>  
             </contentAssets>  
             <!-- This changes for each video -->  
             <customDataItems>  
                 <customData metadataProfileId="22971">  
                     <xmlData>  
                         <metadata>  
                             <METXDistrictDistrictName>Ann Arbor</METXDistrictDistrictName>  
                             <METXDistrictDistrictNum>20</METXDistrictDistrictNum>  
                             <METXSchoolSchoolName>Huron High School</METXSchoolSchoolName>  
                             <METXSchoolSchoolMETXID>33</METXSchoolSchoolMETXID>  
                             <!-- Kaltura wants the date to be in xs:long format, that is, seconds from the epoch -->  
                             <METXVideoSubmissionCaptureDate>1344440267</METXVideoSubmissionCaptureDate>  
                         </metadata>  
                     </xmlData>  
                 </customData>  
             </customDataItems>  
         </item>  
     </channel>  
 </mrss>  

And that's it.

Friday, September 14, 2012

Setting up Kaltura - part IV

We have been working on getting a Kaltura Drop Folder set up.  A Drop Folder is a mechanism where an organization spools content to be ingested in a fixed location, and Kaltura polls the location, watching for content to ingest.

In our case it has taken about a month to get the Drop Folder configured, and much of this delay is preventable if you avoid the same pitfalls we did.  So in the spirit of giving back to the community, here are seven things to know when setting up a Drop Folder.

  1. Host the Drop Folder yourself, do not host it at Kaltura.
  2. Set up an account and a password on the machine, and share them with Kaltura.  To keep things very simple I created an account called 'kaltura'.
  3. Create a subdirectory under the kaltura user's home directory that will actually contain the content to be ingested.  To keep things very simple I used the name 'dropfolder'
  4. Make sure that the kaltura user owns the Drop Folder directory, and that its access controls grant appropriate rights to other users that may need to ingest content
  5. Tell Kaltura the name of the machine.  To keep things very simple I created a DNS CNAME record, kaltura.icpsr.umich.edu, that points to the right machine.
  6. Be sure you have ssh installed and running on port 22 on the machine.  If you normally do not run ssh on port 22 (we don't), don't forget to open a hole in your firewall so that Kaltura machines can reach the Drop Folder.
  7. Tell Kaltura to use sftp and port 22 to connect to your Drop Folder.  Do not try to use a port other than 22.
To re-cap the values we used:
  • Host: kaltura.icpsr.umich.edu
  • Protocol: sftp 
  • Port: TCP 22
  • Login: kaltura
  • Password: XXXXXXXX
  • Drop folder: dropfolder 
  • Drop folder UID:GID:  kaltura:met  
  • Drop folder mode: 2775
(We have a big video project called MET, and automated jobs running with the 'met' GID will need write access to the drop folder.)

You may be tempted to suggest using ssh keys or non-standard ports for ssh.  Fight those temptations.

Kaltura will offer to auto-delete content once it has been ingested.  Accept that offer.

Know that when you delete items from the Drop Folder status window in your Kaltura KMC it will also delete them from the Drop Folder.  This is not obvious, but turns out to be useful.

Now all you need are automated jobs that place content and Kaltura-style Media RSS XML into the Drop Folder.  Kaltura has some nice examples on-line, but they are somewhat trivial.  We'll post some more complex, real-world examples next week.

Wednesday, August 29, 2012

Setting up Kaltura - part III

Some good news and some bad news today.

The good news is that I've tested XML Ingest via the Drop Folder feature, and it seems to work very well.  I was able to upload two videos, add their extensive metadata, and get it working well after fixing a simple (but dumb) typo I made in the XML.  In terms of creating the right XML content, we are in great shape.

The bad news is that I've run into a couple of snags with the Kaltura Software as a Service (SaaS) offering.

The first is actually with the Drop Folder service.  The current solution we are using is where Kaltura hosts the Drop Folder and we use sftp with a password to transfer a bundle of content - Kaltura-customized Media RSS XML plus a pair of video files.  However, what we really need is a locally hosted Drop Folder (which costs less to operate) and a way for Kaltura to fetch the content.  So far Kaltura hasn't been able to make this work.  We had hoped that they could use an ssh key-pair (private at Kaltura and public installed in $HOME/.ssh/authorized_keys) for access, but at this time, Kaltura does not support ssh keys.  So we are kind of stuck waiting for ssh key-pair support to appear, or we are stuck uploading content via sftp and typing passwords (i.e., a manual solution).  Ick.

The second issue is around counting bits.  Kaltura charges by the bit - storing them and streaming them.  That makes a lot of sense, and is actually fine by us.  However......  so far we are not seeing reports or analytics that report usage by the bit, only by the minute.

In a case where one has a single collection with a single delivery platform operated by a single administrative unit, this may work quite well.  The monthly bill from Kaltura goes to one place, and the analytics are very useful for reviewing what's getting played, how often, how much, etc.

But in a case where one has a single collection (like Measures of Effective Teaching) with multiple delivery platforms operated by multiple administrative units, this will prove problematic.  In this scenario we really need a bill that breaks out usage like this:

  • Bits stored for the month - XX GB
  • Bits streamed by delivery platform A - XX GB
  • Bits streamed by delivery platform B - XX GB
  • Bits streamed by delivery platform X - XX GB

Then the partners could carve up responsibility for the bill.  For example. maybe ICPSR pays for ALL of the storage and for the bits streamed by its delivery platform, the School of Education pay for the bits streamed by its delivery platform, and a future partner-to-be-named pays for the bits it streams.

We had hoped to start using Kaltura in September, but my sense is that the Drop Folder issue will push this back.  But the issue around counting is the big one, and I don't have a sense yet for whether this is easy to address or very difficult to address.

Friday, August 17, 2012

Setting up Kaltura - part II

I mentioned in the last Kaltura post that we've set up a Custom Data schema to hold the descriptive metadata for our video content.  Setting up that schema took some considerable effort, and I thought I might share some of the details with this post. For context we are using the newly release Falcon edition of the Kaltura Management Console (KMC) web application.

One creates a Custom Data schema by using the Settings menu in the KMC, and then by navigating to the Custom Data tab.  A button allows one to Add New Schema, and I happened to give ours the name of "MET Extension" since that is the name of the project generating the video.  I did not set a System Name for the schema, and while Kaltura said that was required, the KMC does not enforce it, and lack of a System Name has not yet proven to be a problem.

One adds fields/elements to the schema one at a time using a pop-up window, Edit Entries Metadata Schema.  This can be seriously laborious if you have a large schema like mine with 40 elements.  Lots of cutting and pasting.  Kaltura allows one to export the schema as an XSD XML file, but one cannot import such a file to create or update the schema.

Elements can be Text, Date, Text Select List, or Entry-id List.  Each can be single or multi-value, and each can be indexed for search (or not).  The KMC allows one to supply both a short and longer description for each element.

Text fields are exactly what you would expect, and Text Select Lists are basically pick-lists.  The Entry-id field is useful if you want to store the ID of an extant Kaltura object.  Date is a little tricky since the format one uses to supply a value for this field works one way in the KMC interface - conventional calendar format or pick from a calendar widget - but a very different way when ingesting via XML where it expects an integer number of seconds since the epoch.

We will be ingesting tens of thousands of videos into Kaltura, and so we will NOT be using the KMC to upload videos and to compose metadata.  Instead we will be using their Drop Folder mechanism where one puts XML metadata (including pointers to video files) in a special-purpose local location to which Kaltura has access.  Preparing the XML content that includes both the descriptive and technical metadata is our current project, and I'll report on that process - and the Drop Folder process - next.

Monday, August 6, 2012

Setting up Kaltura - part I

I've mentioned in previous posts that the University of Michigan is implementing Kaltura as its video content management solution.  Kaltura is an open-source video platform that one can install and operate locally, and is also available in a software as a service (SaaS) version.  The U-M is making use of the SaaS edition, and ICPSR is one of three pilot testers.

The off-the-shelf web application provided by Kaltura to manage content, collect analytics, publish content, create custom players, set access controls, etc. is called a Kaltura Management Console (KMC, for short).  A major question for any enterprise using Kaltura is:  How many KMCs do we need?  The answer is:  Just enough, and not one more.

It is very difficult to share content between KMCs, and so there is a major incentive to have the smallest number of KMCs, perhaps only a single one.  However, it is also difficult to "hide" content from others who are sharing the same KMC, and that can cause concerns about privacy, access control, and inadvertent use (or mis-use) of content.  In my mind giving someone an account on a KMC is like giving someone root access on a UNIX machine.

The solution we used at the U-M was to deploy two KMCs for now.  One is for two types of content:  video which is generally available to the public, such as promotional materials, and video which is used in courses via our local Sakai implementation, CTools.  We provisioned a second KMC for ICPSR to use for its content, which falls more into the "research data" category.  This content will require signed agreements for access.

Once Kaltura provisioned our KMC I performed a few initial house-keeping chores:

  1. Created accounts (Authorized KMC Users) for my colleagues on the project.  Each has a Publisher Administrator role.  (Administration tab in the Falcon release of the KMC.)
  2. Changed the Default Access Control Profile to require a Kaltura Session (KS) token.  All content managed by this KMC should require a KS token by the player.  (Settings - Access Control tab)
  3. Created a new Access Control Profile (called Open) which does not require KS.  I don't know if I will need this, but want to have a more open profile available.
  4. Changed the Default Transcoding Flavors to (only) "Source."  Our content has already been transcoded, and so we don't need to pay for the time and storage for additional flavors such as HD, Editable, iPad, Mobile, etc.  (Settings - Transcoding Settings)
  5. Created a Custom Data Schema to hold the extensive descriptive metadata that accompanies the content generated in our project (MET Extension).  This step is extraordinarily tedious since it has to be done field-by-field through a web GUI.  I can download a copy of the schema I created in XSD format; wish I could upload one to create it.  (Settings - Custom Data)
  6. Created a slightly customized player for use with our content.  Wanted to size it to fit our content, remove the Download button, etc.  This is super easy.  (Studio tab)
  7. Created a Category which we will use to "tag" our content.  (In this case I created one called MET-Ext.)  This is mostly useful for searching and browsing within the KMC interface.  (Content - Categories tab)
  8. Uploaded a few videos and set a few of the metadata fields.  (Content - Entries)
  9. Put in a request to our account manager to enable a locally hosted Drop Folder.  This is a mechanism whereby we create a local "fetch" location where an automated Kaltura job can pull content and ingest it.  While one would think that this is a common mechanism for submitting content, the process is slow and poorly documented, and cannot be managed via the KMC.  I'll post more details about the process once I have a working Drop Folder in place.
  10. Created the local infrastructure for the Drop Folder which is really just identifying a machine to play host, and then creating an account Kaltura can use.

These steps got us to the point where we could start putting Media RSS files containing the metadata and pointers to the video into our Drop Folder for ingest into Kaltura.




Wednesday, June 27, 2012

Video at ICPSR - OAIS and Access

We're taking a pretty close look at Kaltura as the access platform for a video collection we are ingesting. Here's why....

If we look at the Open Archival Information System (OAIS) lifecycle, most of the Ingest work is taking place outside of ICPSR. (In fact, other than providing much of the basic IT resources, like disk storage, our role is very small in this part of the lifecycle.) Managing the content and keeping copies in Archival Storage is a good fit for ICPSR's strengths; the content is in MP4 format and has metadata marked up in Media RSS XML, so that's relatively solid.

The big questions for us are all on the Access side of OAIS. Questions like:

  • How many of the 20k videos will be viewed on a routine basis?  Or ever?
  • How many people will want to view videos simultaneously?
  • Will viewers be connected to high-speed networks that can stream even high-def video effortlessly, or will most of the clientele be located on broadband connections?  Is adaptive streaming important?
  • Will support for IOS devices - which do not tend to do well with Flash-based video players - be important?
  • Can people comment on videos?  Share them?  Clip them?  Share the clips?
I have a requirement from one of our partners to build enough capacity to stream a pair of videos - these are classroom observations and each includes a blackboard video and a classroom video - for up to 1000 simultaneous viewers.  That's 2000 videos at a bit-rate (roughly) of 800Kb/s.  So maybe about 1.6Gb/s of total bandwidth required at peak.

And I have the same requirement from one of our other partners who is serving a separate audience.  So that is a total of 3.2Gb/s.  That is a big pipe by ICPSR standards.  (Our entire building that we share with others has only a single Gb/s connection to the U-M campus network.)

If we try to build this ourselves we need a pretty big machine with lots of fast disk (20TB+) and lots of memory and lots of network bandwidth. And if we build it too small, the service will be awful, and if we build it too big, we will waste a lot of money and time.

So a cloud solution that can scale up and down easily is looking pretty good as an Access platform.

Next post:  Why Kaltura?

Monday, June 25, 2012

Video and ICPSR

I've posted a few times about a large collection of video that ICPSR will be preserving and disseminating as part of a grant from the Bill and Melinda Gates Foundation.  I'll devote some time this week to a couple of detailed posts about what we're doing, but one vendor that I'd like to mention briefly today is Kaltura.

Kaltura is a video content management and delivery service that offers both a hosted and on-premise solution.  The University of Michigan is entering into a relationship with Kaltura, and I'm serving on a committee which is helping shape that relationship.  (More on this later.)

I have early access to Kaltura's hosted solution for video content, and I've used that access to upload a few pieces of public domain content plus some minimal metadata.  I then have used Kaltura's tools to assemble combinations of video collections (playlists) and video players, mixing and matching liberally to get a sense for what is possible.

Here's what I have so far:
More on "video @ ICPSR" later this week.

Wednesday, May 23, 2012

How many bits of video will I stream?

We have a copy of video preservation and access projects for the Bill and Melinda Gates Foundation.

One project consists of highly restricted video content, and we believe the demand will be low enough - dozens or fewer of simultaneous video consumers - that we can stream the content quite comfortably from ICPSR.  (ICPSR shares a 1 Gb/s network pipe with one of the other centers at ISR, and the bit-rate of each video is about 700 Kb/s.) A follow-on project consists of less restricted video content that we believe will have broad appeal.  A key question for the IT director is if the demand will be so high that it will exceed our capacity to deliver.

My colleagues are projecting that we will have peak simultaneous usage of 2000 video consumers.  A little back of the envelope math (total consumers x 700Kb/s) makes it clear that our network pipe is too small; we'll need to move the content elsewhere for delivery, or split the load across several network locations to make delivery feasible.  Unfortunately this collection is quite large - 20 TB - and so making lots of copies to spread the delivery across lots of locations will be expensive.

Another approach is to move the content into a content delivery network (CDN).  In this scenario the CDN operator will charge us a fixed rate to store our content and a variable rate to stream our content.  So how much will all this cost?

The storage is easy.  We have 20 TB, and so we can calculate the storage costs quite easily.The streaming costs are more tricky, however.  Typically one's costs are tied to the total number of bits streamed each month, but our only data point is the maximum number of total simultaneous video consumers.  So how do we calculate the expected cost?

We've been struggling with this for a while, and I don't know that we've hit upon a good solution.  But we do have A solution.  Here it is....

What if we were to graph the number of concurrent video consumers?  And what if we assume that the graph will be a curve, a Gaussian curve in particular?

Source: NIST
Our Y-axis can measure the total number of simultaneous video consumers at a given point in time.  We have our maximum height value (2000) as one data point.

Our X-axis can measure time of day where each point is a single second in a 24-hour period.  And we'll choose the starting point and ending point so that the maximum height falls in the exact middle of the graph.

If we calculate the area under the curve this will tell us the total number of consumer-seconds, and we can then multiply that by 700 Kb/s to calculate the total number of kilobits streamed in a 24-hour period.  And we can divide by 8 x 1024 x 1024 if we want to turn kilobits into gigabytes, a standard unit of measurement for calculating streaming costs.

To calculate the area under the curve we need to know the maximum height (2000) and we need to estimate how "fat" or "thin" our curve will be.  (This is related to standard deviation in a normal distribution.)  So if our X-axis is seconds, we might pick something like 60 (for a very pointed curve) or 3600 (for a flatter curve). And if we call the height 'a' and the width 'c' our formula for measuring the area (bits) is:

a x c x SQRT ( 2 x PI )

We can then use fixed rates to turn number of consumers into number of GBs.  And if we have a per-GB price, we can turn that into a total daily cost.  I made a little calculator at Zoho to help with this.  (Note that you must be sure to use the Tab key to move through the form.  Hitting the Enter key or using the Submit button stores the information in a throw-away table at Zoho and clears the form. )

For example, if I think I'll have a maximum of 2000 simultaneous consumers (a = 2000), and I think my curve will be medium width (c = 1800), and my video is 700 Kb/s, and my price to stream is $0.25/GB, then my daily cost will be approx $188.

Friday, October 7, 2011

ICPSR wins grant from the Bill and Melinda Gates Foundation

ICPSR and partners at the University of Michigan received a grant from the Bill and Melinda Gates Foundation recently.  You can find the official link here.

 The link says that the grant is to house and make available to qualified researchers the data collected by the Measures of Effective Teaching project, and than means that you'll soon see yet another ICPSR operated web portal.

In addition to building and operating the portal, the project also requires us to process, preserve, and deliver quantitative data related to the Measures of Effective Teaching (MET) project, which is right in ICPSR's wheelhouse, of course.  The really new element for us, though, is the collection of video and "artifacts" related to the video.

ICPSR will use its existing Restricted Contract System (RCS) to screen applicants who want to access the video collection.  If approved the applicant will be able to access a video streaming server to view the videos in the collection.  The applicant will also be able to access the quantitative data in our Virtual Data Enclave (VDE).

The video collection is large compared to our current holdings of survey and government data.  My sense is that our collection will pretty much double in size, approaching 20TB total.  That's very big for us, but not nearly as big as some collections of video.