Showing posts with label XML. Show all posts
Showing posts with label XML. Show all posts

Friday, September 26, 2014

Similar needs, different results

I recently had the pleasure of concluding two consulting assignments. Though both clients have a lot in common, the outcome was quite different. Why?

Situation and solutions

Both clients are publishers that serve national and international markets. Both offer multilingual products to professional users in PDF and XML-derived formats (such as HTML or ePub). Both need to update their publishing processes to keep up with changing stakeholder demands and are faces with outdated publishing infrastructures.


Image by Sean MacEntee from Flickr

But at one client we advised an Alfresco based solution, whereas the other client we advised a MarkLogic based solution. Being so similar, why two different outcomes?

XML as design choice

Needing editing by countless authors and editors worldwide, client A implicitly chose to set an open and well maintainable storage format as a standard. Generally, (X)HTML will not suffice for publishers, failing at standardisation of contents and the quality of the printed products. So XML it is.

Some excellent Commercial-off-the-Shelf XML products are available (including SDLs Live Content and Alfresco based Componize), but the of-the-shelf part usually means complying with their internal standards: DITA. Although DITA is a great XML format, it is not suitable in all cases, as it was in this case.

Wanting more specialised workflows and interfaces, we concluded that a custom XML based CMS combined with an online XML editor would fit this client’s needs best, with MarkLogic being the best platform to implement this. Binary content (like PDF) will be supported in the system through a layer of XML Metadata.

XML as an option

Client B’s content is mainly created in separate and formalised processes, with lots of content being available only in PDF. Editing being out of this equation, re-use and workflow were of much greater importance.

This client also sees the opportunities for XML and more dynamic publishing, but the hurdles are substantially bigger, and the rewards in editing would not be felt. Rather than going all-in on XML, client B will transform selected publications to XML.

For this client the out-of-the-box functionalities like workflow and flexible content storage tips the scale to an Enterprise Content Management System. Alfresco being the ECMS of choice, this strategic platform will provide the room to manage content and changing customer needs that will characterize enterprise IT for years to come.

Conclusion

It is often said that the devil is in the details. But these details often drive big design choices. Decisive product-ownership provides the roadmap for such strategic choices.

Monday, May 5, 2014

Big Content challenges

At Dayon, we are used to work with Big Data. Coming from a publisher’s background, we have provided content solutions to publishers since 1997.

I read some stories about Big Content, and was intrigued that Gartner saw Big Content as the unstructured part of Big Data. To me, Big Content is the structured version of Big Data.

Let me explain this and address some challenges and Big Content technologies.

Planned Variety

In Terms of the three Big Data V’s (Volume, Velocity and Variety), publishers content is odd. Since the goal of publishers is to make a profit from providing content, content must be able to be published to a vast arrange of channels. To enable this, content must be structured (preferably in XML) and enriched with metadata. Any Variety is planned, because unplanned Variety leads to unplanned structures and/or unplanned publications.

Data is generated, where Content is handcrafted. Tweets en Facebook-posts are only lightly structured, but Blog posts are already quite structured. Some numbers by Chartbeat can be found here and a useful insight by Fastcompany on the rise of “Big Content” as a marketing Tool.

Publishers Content is usually completely structured: XML + Meta Data, sometimes already as RDF Triples (read my earlier Blog post on Semantic Technologies).

So to me, Content is structured Data. Big Content problems differ from other Big Data problems, where handling the Variety to understand your data is a big issue. Therefore, I would like to label the publishers challenge to be a Big Content challenge.

So how big is Big Content?

A quick scan at some of our publishing clients provided these numbers (XML only!):

  1. Publisher 1: 10 million files, 25 GB
  2. Publisher 2: 750.000  files, 15 GB
  3. Publisher 3: 150 million files, 15 GB
  4. Publisher 4: 1 million files, 15.000 new files per day (max)
  5. Publisher 5: 45 million files, 20.000 new files per day (max)
  6. Publisher 6: 500.000  files

Challenges of Big Content

With these numbers in mind, what are the challenges for Big Content?
  1. Volume - XML: Are 30 million XML files a challenge? Or 25 GB in XML a challenge? It really should not be, but in reality I have met quite some technologies struggling with these amounts. An XML system should be true XML to handle this amount of data. XML isn't hard. Doing XML right is hard. If you don’t do XML right, 100.000 files or 1 GB of XML can get you plenty of headaches.
  2. Volume - Other file types: Alas, not all Content is XML. Many Publishers still manage huge amounts of HTML, PDF or other file formats. With PDF, huge numbers often also turn into huge volumes because multi-channel and hence print-quality PDF is stored.If you have to index lots of other file types, do a proper intake process per file and weed out the corrupt and the largest files.
  3. Volume - Subscriptions: At various clients I encountered the problem that Big Content is offered in large amounts of different Subscriptions. Whereas a large amounts of different Subscriptions are not a problem in itself, the combination of Big Content and Big (number of) Subscriptions often is. So if you offer lots of data, be smart about the number of Subscriptions.
  4. Volume - Triples: Nearly all Publishers storing Big Content are looking into Triples as a way to store and link Meta Data from their XML files. Storing your Meta Data in a Triple Store, and Linking it to the Linked Open Data can be a very good idea, but this calls for a Big Triple Store. A set of 1 billion Triples isn't exceptional, but also requires Big Content Technology.
  5. Velocity - Real Time Indexing: Failing at real time indexing is usually the first sign that you are becoming a Big Content publisher. Many technologies struggle with incremental updates, needing complete re-indexing, which in term leads to strange solutions such as overnight indexing, flip-flopping or indexes out of sync with the rest of the front-end.
  6. Velocity - Real Time Alerting: The value of Content depends on its relevance, and timeliness is a huge factor in relevance. Real Time Alerting will offer a competitive edge to content users. To provide Real Time Alerting, XML store need to handle alerting efficiently (using minimal resources) at load time
  7. Variety - Presentation: A Big Content challenge can be how to present all of this Content. If a simple “What’s New” view results in 20.000 hits, what are you going to show the customer?The most used solutions are:
    1. Provide a Search Only interface
    2. Provide as much structure from Meta Data as possible to assist the user in drilling down to the most useful Content
  8. Variety - Enrichment: If the Meta Data you need to provide useful segmentation of your Big Content to your end users just isn't there; there is a need for additional Enrichment. Big Content will (due to costs) call for automated enrichment using Natural Language Processing

Big Content Technology

At Dayon / HintTech we strongly believe that Big Content challenges require specialized Big Content Technology. Here are some of the Big Content Technologies we have implemented:
  1. MarkLogic
    Several of our Big Content clients have selected MarkLogic as their content platform. I believe that MarkLogic is the best XML store and indexer available at this moment.
    As a big bonus, MarkLogic comes with all kinds of useful features such as XQuery, an Application Server, and now even a Triple Store.
    Find out more about MarkLogic at MarkLogic World Amsterdam and meet us there!
  2. OWLIM
    In our project at Newz we needed a Big Content Triple Store. We found OWLIM by Ontotext to provide an excellent Big Triple Store, as did BBC and Press Association.
    W3C maintains a list of Big Triple Stores, with BigOWLIM as one of the top products.
    We also selected Ontotext as our partner for their Semantic Tagging capabilities.
  3. SOLR
    We have also implemented SOLR for Big Content collections. SOLR will not face all of the Big Content challenges, but is a great Open Source search engine.

PS: After writing this blog, I feel like renaming Meta Data to Meta Content. Probably better if I don’t…

Thursday, March 20, 2014

Top 4 Web based XML Editors

Why web based XML editing?

Publishing natively from XML still is the best solution for publishing demands. Not only if you want to publish to print. Products will be more functional if they are formatted on basis of their meaning, and not their HTML appearance.

For years, XML editing was cumbersome. Costly, non-intuitive XML editors drove all but the most technical editors crazy, and induced Tag-Terror in others. Plus, they needed to be installed on your local system.
Many developers have tried to use Word as an XML editor, but only the bravest have succeeded. Applying schemas after a submit, or translating Schemas to Templates all resulted in much manual correction or unsatisfied editors.

The dream remained of a Web based XML Editor with the following qualifications:

  • WYSIWIG editing
  • No local installation
  • Can be used by many authors 
  • Minimal training necessary
  • Guaranteed schema compliance

Xopus was one of the first Bracket-free online XML Editors, but there are more contenders now. A concise overview of the top web based XML Editors (and not just DITA!):


Xopus

Xopus was one of the first real web based XML Editors. It received huge attention. Eventually early adapters noticed that it still enforced XML rules. It did not provide MS Word-like editing – as can be expected, but it was clearly a big step.
Xopus was acquired by SDL last year, probably to support the new LiveContent product. Version 4.3 of Xopus was launched in February 2014.


FontoXML

The second half of 2013, FontoXML was launched. FontoXML offers a very intuitive interface and some really nice real-time features.


Xeditor

In 2013, Appsoft launched Xeditor. Also a very user friendly web based XML Editor.




<oXygen/>

<oXygen/> is known for its XML products, but it has also seen the need for online editing. An Author Demo Applet for editing DITA can be found on the <oXygen/> website.


You choose!

Which XML Editor suits your needs best depends on your specific editing needs, but I am happy that there are multiple vendors now wanting to suit your XML Editing needs.