Showing posts with label Content Management. Show all posts
Showing posts with label Content Management. Show all posts

Thursday, January 14, 2010

Delivering Integrated Information from Data Silos Using MODS

Lyn Robinson of the Burton Group wrote an excellent paper describing a "Methodology for Overcoming Data Silos" (MODS).; a "groundbreaking project structure" for bridging silos and delivering integrated information from decentralized, disparate information systems.

Interesting concept. A lot of thought, time and money has gone into breaking down information silo's over the past few years. Robinson recognizes that this is frequently unsuccessful or, at best temporary, but that the need for integrated information remains very important - even mission critical. So instead of re-engineering the underlying information systems into a single integrated replacement, MODS proposes to implement an information bridge that fills the need for integrated information from across these silos.

So how does it work? "Each MODS project results in the definition of a few relevant and important enterprise data types, as well as a registry that tracks and reconciles the instances of those data types in the disparate silos throughout the enterprise. This registry functions somewhat like the old card catalogs for books in a library. Using a MODS registry, businesspeople can discover where to find particular types as well as instances of relevant data within the disparate silos scattered all over the enterprise." And, it's supposed to be cheap (well, cheap in comparison to development projects).

Wow. Cataloging data - maybe a librarian should be called in! Robinson describes a typical MODS pattern as "establishing a registry or system of record for a particular set of closely related enterprise data types. The scope of these data types should be determined by a conceptual data model, which defines the needs the business has for information regarding a particular business topic. The best MODS projects are narrow in scope, so the more specific the business topic, the better". Here are the steps in the pattern:

  1. Model the relevant data.
  2. Clean and Understand the Relevant Data
  3. Classify
  4. Convey
  5. Control
As I read this approach, I got to thinking, this is a possible application best served by a TopicMap. After all, TopicMaps were originally developed for indexes or catalogs. It seems like a natural fit. TopicMaps can implement many of the features MODS calls for: data typing, data relationships and styled queries. And I also began to think about social networking features like folksonomies and how these might apply to MODS.

A lot of my personal focus is related to content management and publishing. As I consider this, Content Silos are just as prevalent in publishing companies as other information type silos are in other organizations. So MODS may well have a place in content management.

So, in the next several blog posts I'm going to explore TopicsMaps and perhaps Folksonomies as technology solutions for implementing a MODS project for content in publishing.

Stay tuned...

Tuesday, October 6, 2009

RSuite User Conference - 10/6/2009

I'm attending the 3rd RSuite users conference in Philly today, by invitation of Really Strategies' Lisa Bos. I've been tracking RSuite since its inception a few years ago - the idea of a decent CMS based on Marklogic always hit me as a good niche product and by all indicators, RSuite is a success, especially in light of how many CMS offerings there are in the world.

I'm here for the technical track - to learn more about how to write customizations for RSuite - I'm a developer afterall. But the Keynote from Howard Ratner (CTO, Nature) is an excellent intro to what comes after CMS.

Nature is using RSuite as an integration tool between various pub tools and the Marklogic content repository. In the Keynote, Howard makes excellent points about the need for content classification, which (of course) is why I'm writing about it! Nature is using a mixture of automated content recognition tools and editorial tools. Also mentioned open search technology - search widgets that create a viral access to the Nature content. One very interesting thing Howard talks about it the "Work Arounds" that are developed because either the content or the technologies don't do the exact job needed. What is the cost of these?

Lisa Bos then gave us a demo of the current version of RSuite. Nice clean and simple to use browser-based GUI. New features include Digital Asset Management, Workflow Configuration, editing (in oXygen and XOpus - one of my fav XML editors). But, IMHO, where's the beef? What does this have that free/open-source solutions (like Alfresco and/or eXist) don't? Or, for that matter, our own Tractare? Other new features - search enhancements, reporting, GUI customizations, better API and the ever needed performance improvements.

Friday, September 19, 2008

Semantic-Content Management

We've been developing a tool set, or framework, or whatever for several years. We call it "Tractare" which is Latin for "to handle, manage, perform". I'm a sucker for that kind of naming.

Anyway, what's it all about? Well, as the Internet becomes more saturated with raw information, keyword search engines really aren't enough to locate that needle in the haystack. We (content providers) need to describe our content in a way that users "get" -- we need to describe the "aboutness" of our content.

An example I like to use when speaking on this subject is this: If you were to use google to search for "retarded", you would get a gazillion hits. But few, if any of those hits would have come up using the politically correct phrase "intellectually challenged". This is because a keyword engine like google depends on the actual presence of the keyword, either as text in the content or as metadata. Now, you could encode both forms of this concept as meta-data on your web-page and it would be found. Now, if you are a psychologist or someone working in mental health, you'd probably be getting the results you want. But if you are a firefighter, the word "retarded" has a whole different meaning. How do we express that? The answer is in several parts of course. But first, we need to capture the meaning of the content; the "aboutness". We need to associate the "firefighter" concept with the content that pertains to fighting fires. This is what Tractare permits us to do.

Tracare is a framework. It's not an off-the-shelf product. It is built on the idea of topic maps -- organizing content around indexes and concepts. It's true power lies in a combination of searching and navigation tools that allow the user to narrow the scope of their work to a set of concepts. We build custom CMS and delivery solutions on top of it.

The CMS systems we build usually include features found in social networking, including folksonomies (as well as traditional taxonomy and classification support) and ranking/commenting. These features allow content providers to apply semantics to content in a number of new and different ways.

The delivery systems we build often include a number of search and navigation interfaces that web users have come to love, including mashups, classification searching, semantic browsing and so on.

Friday, March 28, 2008

Using Folksonomies in Content

I've been thinking more and more about folksonomies as a replacement for traditional taxonomies. We can create all the tools we need, or work off of existing tools such as del.icio.us. But in the end, size does matter. For a folksonomy to work, a *lot* of people have to look at and tag the content. To me, this means content has to be exposed to readers, and lots of them, to tag the content.

Traditionally, the classification operation has been something done under control and part of back-room content management, by a small, select group of indexers. So first, we have to decide to relinquish some control. While this may seem scary, in reality the volume of tagging makes up for the lack of specific control. And, we can build the tagging tools so that editors and indexers can follow behind the tagging and clean it up. But if we can achieve a large volume of tagging, the volume and repetitive nature of the tagging will create common tags, threads and relationships.

Second, we need to find a large group of taggers. Depending on the nature of the content, this can be accomplished in a couple of ways. First, and perhaps easiest, is to expose content to the web. Perhaps through incentives, taggers can be enticed to tag. And, of course, staff of the publisher should be encouraged to participate as well. Failing that, a publisher could look at a human-automation engine such as Amazon's Mechanical Turk, where large numbers of minuscule tasks that are best done by people are spread out over many people for a small fee.

Monday, March 19, 2007

Content Management Frameworks -- Tractare and Configuration

There are a large number of Content Management "Frameworks" out there -- http://en.wikipedia.org/wiki/Content_management_framework lists just a few. What is a "Content Management "Framework?" In essence, it is a programmable API for creating customized content management systems. In other words, a programmers way to create (using the framework) a CMS that supports all the stages of content lifecycle: Organization - Workflow - Creation - Repository - Versioning - Publishing - Archives (or some subset thereof). In essence, a CM framework is a foundation on which a custom CMS can be built.

The key features I would expect to see in a CM framework include (http://www.cmsreview.com/Features/Lists.html):

  1. a way to acquire both text and non-text content (acquisition, aggregation, authoring)
  2. a way to store and retrieve content
  3. a way to control workflow (roles/permissions, checkin/checkout, messaging/routing)
  4. a way to control versioning
  5. a way to control personalization and localization
  6. an interface to administration (reporting, management, etc.)
  7. a way to control content delivery (extraction, slicing, publishing, syndication, update)
  8. a way to implement business rules
  9. and others...
The trick is making this work as a framework. How much needs to be customized by a programmer and how much can be done by a very tech-savy non-programmer.

In our CM Framework (Tractare), we permit most of these features to be customized by scripting. For example, many organizations have IT departments that mandate the use of a particular DBMS (Oracle or SQL Server, for example). So in Tractare, there is a fairly simple configuration setting that allows selection of the database interface driver. But, since Tractare is a CM framework, we also expose the driver interface so that a programmer can create a driver for a database for which we have not supplied one. We make heavy use of configuration files so that non-programmers can customize the CMS.

Another area that we use scripting is in the interface elements (the "view" of the CMS). Internally, Tractare generates most interaction responses as an XML stream. This way, the entire user interface to Tractare can be tailored using XSLT scripts.

Depending on the application however, Tractare could still require an investment in programming. It is written in JAVA and intended to run as a web application (using servlets, etc.). So extending the code is a matter for a JAVA web programmer. But many features can be customized without this level of involvement.

Sunday, March 11, 2007

Open Publish Conference

I got to speak at the Open Publish conference in Baltimore (http://www.open-conferences.com/baltimore/) on Friday, 3/9/2007 amongst many of the content management and publishing luminaries. It was humbling. While the conference was small, it was very interactive and the audience was firing questions both during and after the talk. Kept me on my toes!

My particular talk was on topic maps, and on using topic maps as an information architecture to improve hypertext linking. The premise is that users frequently click on links embedded in web pages that take them somewhere they didn't want to go and that with better knowledge captured on the "aboutness" of the link, we (as content providers) can give the user more information and options about the links they are following (anyone interested can get the presentation slides from the conference directly or contact me and I'll email them).

My presentation was either fabulous or boring (depending on whether you're asking me or an attendee). But what I came away with were a bunch of questions about how to glean a topic set from extant content. I've spoken on the subject of topic maps on several occasions and many of the questions follow this same theme -- "How do we collect and organize the topics of the content in a meaningful way?"

I always have some lame answer -- "that's a job for the subject matter experts." While this is true -- the current state of the art is somewhat limited in our ability to parse content and determine subject matter, it really begs the question. Often (in my experience) there is something with which to begin this process. Usually, our content has already been touched in this way, from something as simple as the application of keywords to creation and maintenance of an outside topical index and the positioning of a particular content object within that. And often there are tables of content that can be used to further glean the "aboutness" of the content. What we (developers) need to consider is how to create a toolkit that both captures as much of this meta data as possible and imputes a relationship structure to it that in the least can serve as the foundation for a subject matter expert's work, and in the larger sense can provide a fully automated creation of a functional, integrated taxonomy.

At this conference, I was asked if I knew of any translation program that would convert the output of an index creation/management program (CINDEX) into a topic map -- a good example of what I'm talking about here.

I'm going to explore this further. We (Retrieval Systems) have a bunch of content processing tools, including some tools to work with the CINDEX data formats. Perhaps we can develop a reasonable toolkit for this kind of conversion.