Wednesday, February 22, 2012

Metadata: the answer to the digital dilemma


For one of our clients (a national broadcast news media organization), we manage 2.2 petabytes of digital assets, which is growing at a rate of roughly 200 terabytes monthly.  At this rate we are eclipsing anything imagined 10 years ago, due to technology advances, high-definition, and the speed at which information is available and shared.  Storage companies like EMC are preparing for even faster growth speculating that we will see 35 trillion additional gigabytes of data added to the world before 2020….or 35 zettabytes (a word that didn't exist prior to 1991…and forward thinking even then). 

Handling trillions of bytes each month, whether creating, meta-tagging, consolidating or replicating, means that we are part of the above digital dilemma – what to do with all the data.  Our archives are bursting at the seams with physical assets that have been digitized, backed up and protected…so now we are paying for not only the digital storage, but the physical storage for an asset to live, just in case the digital asset fails (and its backup fails).

In the physical world we have organized systems, like LC, Dewey, stacks, cards, etc…that all have the meta-data" about the physical asset.  If all else fails, the asset itself has enough information about it to re-create the meta-data; however, in the digital world that isn't always the case.  Date, author, photographer, videographer, producer, actor, etc.  ….these are all things that can be lost forever if the digital asset isn't properly meta-tagged when originally digitized.  Too many skip or skimp on this critical step.  Saving a buck on the metadata can cost you thousands later.  Lost productivity hours are just the tip – if you don't have the right meta-data on the digital asset you risk losing it for good …in the heap of trillions and trillions of bytes that are spread across trillions of DATs loaded in millions of SANs globally.

Even if you are only talking about a single organization, the numbers are staggering.  Unstructured data is likely to become the largest single expense for businesses, even surpassing staff, within the next 10 years.  In 2002 the total amount of information created in the world was 5 exabytes.  By 2006 that number was 160 exabytes.  Today facebook.com is the size of the entire internet in 2004 according to Geohive and Facebook.  Youtube estimates that 35 hours of video are being uploaded to the site every minute.  Pingdom estimates that there are roughly 300 billion emails sent daily!  According to EMC, the world’s information is doubling every two years. In 2011 the world created a staggering 1.8 zettabytes. By 2020 they estimate that the world will generate 50 times the amount of information and 75 times the number of "information containers" while IT staff to manage it will grow less than 1.5 times. This means that properly meta-tagged digital assets will be critical to successfully manage your digital assets.  New "information taming" technologies such as de-duplication, compression, and analysis tools are driving down the cost of creating, capturing, managing, and storing information to one-sixth the cost in 2011 in comparison to 2005; however, even with proper reduplication and "taming" the internet growth is speeding up, not slowing down.  So what's bigger than a zettabyte? A yottabyte, which is 1000 zettabytes.  While it seems as though that's a long way off, it will be here before we know it and we'll be off to add a new word to the dictionary.  

To be able to process, search, absorb or synthesize that data you must have exceptional metadata.  Without it, think of your video or image asset as a grain of sand at the bottom of the ocean.  You can describe the shape color and size of that grain of sand all you want (after the fact) but the odds of you coming up with the exact grain you were after, is impossible to comprehend (by 2015 this number is estimated to be 1 in 1.25E+22)…and the odds are stacked against you more and more by the second.  In fact, imagine while you are looking on the beach for that grain of sand, 15.6 million beach volleyball courts worth of sand was being added to that beach.  With the right metadata you are able to instantly search for that grain of sand, dive in and come up with the thousands of possible grains that match the description…and then drill down from there to get to your target.  That's the real power and value of what we offer our clients…our staff, working across the world, enable information, putting the power behind the search and allow trillions of bytes to be reviewed, to find that one specific item.

Tuesday, February 14, 2012

Early Industry Adopters: Government Agencies in the Digital Marketplace

When we look at the advantages we gain by migrating our intellectual properties from the physical paper world to the Digital Marketplace, it is hard to find better cost savings, return-on-investment, and increased business operational and end-user efficiencies, than those realized by early adopters: academic institutions, government agencies and law firms.

Our Machinery of Government is comprised of federal, national and state agencies, established for the oversight and administration of specific functions to preserve our history, and our way of life. These government agencies include everything from our national security and justice departments, the Environmental Protection Agency, the US Postal Service, Bureau of ATF&E, National Aeronautics and Space Administration to our Library of Congress, where our national history archives are preserved. The assets created and housed within these agencies could easily be considered among our national treasures.

Its worth mentioning, that LAC Group works with all of the government agencies mentioned above, as a business partner in the Digital Marketplace. The Library of Congress specifically is an important relationship for us, not only as a client, but also as a partner in standards for all digital libraries, as all other industries follow suit.    

Some of the primary standards (using the XML schema language of the World Wide Web Consortium) include:

·      MARCXML and MODS – metadata to describe the content of a digital item
·      METS and MIX - metadata formats for media and environment of a digital item
·      PREMIS - metadata format supporting preservation activities for a digital item
·      SRU - protocol for search and retrieval in the digital environment

All of these standards are maintained in the Network Development and MARC Standards Office of the Library of Congress, and available online now, for 312,000,000 Americans to access, any time, and virtually from anywhere.

Join me next week as I continue my discussion of digitization benefits and standards.

"A good library is a place, a palace where the lofty spirits of all nations and generations meet."
– Samuel Niger

Wednesday, February 8, 2012

Early Industry Adopters: Academic Institutions in the Digital Marketplace

Untitled Document
When we look at the advantages we gain by migrating our intellectual properties from the physical paper world to the Digital Marketplace, it is hard to find better cost savings, return-on-investment, and increased business operational and end-user efficiencies, than those realized by early adopters: academic institutions, government agencies and law firms. 

According to the National Center for Educational Statistics, in 2009, there are 6,632 colleges throughout the United States. Each of these academic institutions generate thousands of new curriculum study books, resources, papers and tests each year, in addition to, growing and maintaining private libraries and art collections. An accredited school can have a core collection of 125,000 volumes and up, in a wide spectrum of different media.

Basic standard metadata content categories provided below for media:
T —Text and illustrated printed material including: books, journals, manuscripts, maps, line art, explanatory tables and drawings.
PR — Visual arts and/or pictorial materials including: prints, photographs, drawings and paintings. 
PT — Photographic negatives and transparencies. 
AR — Specific purpose images produced by reformatting aerial, medical and scientific images, architectural and engineering line drawings and blueprints.
3D — 3-dimensional works of visual art, objects and artifacts located in archives, galleries, and museums

In 2011, an estimated 19.7 million (college) students began their new academic year, creating enormous demand for access to the 816 million books and serial volumes in their libraries, and the 9,221 public libraries, in their communities.
 
Digital archiving has forever changed the world of education, as we knew it. Millions of books and collections have been converted into digital assets, providing online access, in real-time. Digitization has also allowed us to meet today's changing needs, by offering virtual online classrooms, universities and learning centers, and have expanded students of every age, including Nola Ochs - who is the world's oldest college graduate at 95 years old.  

The digitization that has, and is taking place within our 105,338 academic institutions has exponentially increased our educational service offering in America, by providing virtual access to millions of educational materials only available in the traditional physical libraries of our past, to millions, within seconds. We honestly haven’t begun to realize all of the benefits in education yet, as a country, or world even, simply because there are so many, for so many.    

LAC-Group Success Story, developing and implementing (ILS) integrated library management with a 4-year college client.

Join me next week as I discuss digitization benefits and standards within government agencies.

"I made use of the college library by borrowing books other than scientific books, such as all of the plays by George Bernard Shaw, the writing of Edgar Allan Poe. The college library helped me to develop a broader aspect on life." - Linus Pauling, Scientist

Tuesday, January 31, 2012

Document Imaging: With and Without Metadata Tagging

Over the past few weeks, I've talked about how metadata is the “data about our data and/or data containers” and tagging that data is critical for the proper document imaging and archival of our intellectual properties. Without dynamic or “smart” data tags, our document images are available online in their repository, but accessing those documents, or the specific information contained within those document images still presents a huge challenge.

Below I've put together a short commentary on the history of the DDC for the purpose of exploring the accessibility benefits of document imaging with, and without metadata tagging, as seen below:



A Brief History of the Dewey Decimal Classification System (DDC)

The Dewey Decimal Classification System (DDC) marked the beginning of the modern library movement in the nineteenth century. The man responsible for its creation was Melville Louis Kossuth Dewey, known to most today as, Melvil Dewey.

Melvil Dewey created the DDC in 1873, and had it published and patented in 1876.

The DDC system grew from its first edition in 1876, and has been translated into over 30 languages to date, including: Arabic, Chinese, French, Greek, Hebrew, Icelandic, Italian, Korean, Norwegian, Russian, and Spanish.

Over the last 135-years, the Dewey Decimal System has been embraced for metadata standards in over 200,000 libraries, in 135 countries, for the proper classification, aggregation, identification and location of specific books and sundries, all over the world.


Document Image (only): Users who want to access this document online have to know the exact title of the scanned image (The History of the Dewey Decimal Classification System), or select the title, author and/or creation date from an index list of all scanned document titles in order to access it. Users also have to have proper access to the proprietary LAC-Group repository where this document is stored.

Document Image with “smart” metadata tags: By adding metadata tags to key-words (all highlighted in  blue font) including: document title, document author, date created, intellectual property of, and all of the important facts about the topic, this image is converted to a 3-dimensional dynamic data set.  Users can now query the LAC-Group database, or use an Internet browser to access this data in the Digital Marketplace by simply typing in any one of the following words, or query phrases:
-    History of Dewey Decimal System
-    History of DDC
-    White Papers on DDC by Rob Corrao
-    DDC by LAC-Group
-    Melvil Dewey
-    DDC creation date
-    DDC publication date
-    DDC patent date
-    DDC languages
-    No. of libraries using DDC
-    No of countries using DDC
-    Documents created on 1/24/12

In addition, if dynamic hypertext links are included for each of the above key-words, the user not only accesses this scanned image, but potentially dozens more that have metadata tags and links to this subject matter. Metadata tags also provide users with quick access to specific facts within the scanned images like: DDC creator - Melvil Dewey, 1873 creation date, 1876 publication and patent dates, and DDC in Korean, Arabic or Hebrew even, saving immeasurable time in research.

Join me next week as I delve into some of the specific benefits of proper document imaging with metadata tags within specific industries.

"A library's function is to give the public in the quickest and cheapest way information, inspiration, and recreation. If a better way than the book can be found, we should use it." – Melvil Dewey (1851-1931), American Librarian and Educator.

Tuesday, January 17, 2012

What is Metadata Tagging?

Tagging your metadata (“data about your data and/or data containers”) is the key to proper classification of your assets in your archival system, in order to provide your employees, clients and prospects proper online identification and access to those assets. Metadata tags are typically words, images, terms and other identification markers that transform a simple image into a dynamic document. When we examine the benefits of metadata tagging in our personal lives, it’s easier for us to grasp the need, and the enormous benefits for metadata tagging our document images and digital assets within our businesses.

Metadata tagging is already part of our every day lives, helping us all find what we need quickly in the virtual marketplace. Some great examples of how metadata tags save us time and money are easily found in the online applications many of us now use every day, whether Google searches, or in the social and professional networks, Facebook and LinkedIn, YouTube, and Twitter (to name a few). All of these applications include metadata tags for our names, the companies we work for, our network of peers and friends, and our bibliographic information, our authorship of quotes, blogs, articles, books, photography, videos and art. These tags are commonly in the form of dynamic hypertext or web links, Internet book-marks and key-word tags. (Metadata tags, in the form of dynamic links, embedded above for your convenience, allowing us to access all five of these applications easily from this one document.) Metadata tags also enable us to type in a key-word or name into Google or other Internet search engines, and within seconds have an index listing and links to access 100s of documents and sites on that specific subject or person.  Metadata tags have literally changed our research time from days, weeks, perhaps months, to a rapid method for exploring documents and records, within a matter of seconds.

While libraries were already well positioned to convert their assets from the Dewey Decimal System to this dynamic digital environment early on, academic institutions, government agencies, and law firms weren’t far behind them. Today, it’s difficult to find any industry or specific business that hasn’t or won’t benefit from digitization and metadata tagging. Next week, join me as I explore specific examples of document imaging versus dynamic imaging with metadata tagging.

“I basically did all the library research for this book on Google, and it not only saved me enormous amounts of time but actually gave me a much richer offering of research in a shorter time.” – Thomas Friedman

Wednesday, January 11, 2012

What is Metadata?

Metadata has been a business term closely associated with cataloging archived information dating back to 1876, starting with the creation of the standard proprietary library classification system, the Dewey Decimal System.  Remember the old 3x5 card catalog system in your school or community library? Those 3x5 cards contained the book title, author, subject and a short synopsis, and an alphanumeric identification number that provided readers with the section and shelf in the library where the physical book was located. The information displayed on those 3x5 cards is the metadata for the library’s assets. Over the last 135-years, the Dewey Decimal System has been embraced for metadata standards in over 200,000 libraries, in 135 countries, for the proper classification, aggregation, identification and location of specific books and sundries, all over the world. Today, metadata is not only a common term for library assets management, but it’s a term we hear in almost every industry as businesses expand their presence and assets into the virtual marketplace. Specifically, metadata refers to “data about the data”, as well as “data about the containers of data”, also known as the structural metadata.

Most often metadata is an information set that includes:
  1. means of creation of the data
  2. purpose of the data
  3. author of the data
  4. date and time of data creation
  5. placement of the data on a computer network
  6. the standards used or followed
Join me during the next several weeks, as I cover specific industry examples of metadata, metadata tagging, and the differences between document and image metadata, indexing and cataloging to optimize your asset archival and retrieval.  Ways in which organizations are using enhanced metadata to increase sales.  Please respond here, or contact me directly if you would like me to highlight examples and benefits to your specific industry. 

"Books constitute capital. A library book lasts as long as a house: for hundreds of years. It is not, then, an article of mere consumption but fairly of capital, and often in the case of professional men, setting out in life, it is their only capital." – Thomas Jefferson, (1743-1826), 3rd President of the United States.

Tuesday, January 3, 2012

Pros of Outsourcing and Hiring Information Professionals (Part 4 of 4)

LAC Group was founded 25 years ago as Library Associates, to meet the growing demand for outsourced librarians, for physical asset organization and research, in businesses of all industries. Like you, we too had to evolve to meet the changing needs of our clients (hence the name change), and their clients. As technology launched us all into the Digital Age, we responded in kind with an extensive investment in technology and training for our team of librarians, and support for the information management staff of our 1,000+ clients. Today, we have almost 400 information professionals employed worldwide, with varying skill sets, that work in the information management field directly for our clients from inside their organizations. This trend for outsourcing or contracting information management professionals continues to be well received in almost every industry, considering many business owners are still getting their heads around the technology, and the demand by their users for multimedia access to all data in real-time. From a business application perspective, these professional physical and digital asset managers are your information gurus; they are responsible for finding, using, managing and sharing your information, both internally and externally.

Outsourcing some or all information management functions is not only viable, because it alleviates the time, money and resources necessary for hiring and training staff on an entirely new position, with an entirely new skillset, but it also eliminates the additional ongoing operational costs associated with an additional employee including: payroll taxes, health care, general liability and workman’s compensation, to name a few.   Additionally it allows your business to focus on its core business.  Many have the misconception that outsourcing means a reduction in staff or wages – while there are some companies in the industry that have perpetuated this perception, this is not always the case. Hiring an outsourced information professional (in-house) may be not only the most cost effective option for you if you, but a boost to the employee. Further, it allows your organization to develop a strategy for your business for the long run.  Regardless whether or not your information professional is an employee or contracted from a company like LAC Group, the role they play in your organization is significant, and will increase in value over time. We have only begun to scratch the surface of savings ahead for us all, through this environmentally favorable demand for digitization in all commerce across the globe.

Recently, Information Today published an article: Revitalising Outsourcing, thepros and cons of outsourcing for information professional, authored by Iain Dunbar (Oct. 20, 2011), who is LAC Group's GM of UK Operations.  

“With the enormous and steady increase in the volume of our literature, we must rely more and more upon sympathetic selection, judicious editing, and the indexer who knows where to exercise discretion. Any simpleton can write a book, but it requires high skill to make an index.” - Rossiter Johnson