Showing posts with label Data Ecosystems. Show all posts
Showing posts with label Data Ecosystems. Show all posts

Monday, May 11, 2009

Digital Small Business and the Web

A digital economy is far more than simply having lots of internet start-ups. A digital economy is as much about the use of technology as creating technology. The problem with a lot of the conversation around knowledge and digital economy is this requirement is lost as people focus on creating the next Google.
Just as much needs to be done to have small business taking advantage of technology to improve their business as trying to fund new technology companies. It is important to understand the issue here. To have a truly digital economy, small business has to use technology and particular web technology to their greatest potential.

This covers how a business interacts with their customers, with their suppliers/partners and how they manage information internally. You could split these up but to really make a difference small business needs to integrate via workflows (I elaborate on workflows here) not through insiders and outsiders paradigm.

For example, instead of plumbing company having one booking system used by employees to book appointments and another for website bookings rather use a service like BookingBug that both customers and employees can use to make bookings.

But that is only one step. Let’s look at the whole workflow. Basic plumbing workflow is:

  1. appointment is made
  2. plumbing task is done
  3. payment is made.
A customer comes to the site and books an appointment for a plumber. The BookingBug widget shows appointments that have already been booked either by other users on the site or by staff. They pick an appointment and fill in the details. A plumber is sent the information to their iPhone with the details (with the iPhone representing smartphones). The plumber does the job and fills in an invoice via Freshbooks App on the iPhone. This is sent immediately to the customer who can then choose to pay via the payment application on the plumber’s iPhone.

The invoice details are logged to Xero, the accountancy software, which is then followed up with the payment details when payment is made. Any supplies consumed during the plumbing task are sent to the supply management software which logs the usage and re-orders as required. The details of the job and correspondence with the customer are all logged in 37Signals' Highrise.

By bringing in effective web technology (be it SaaS, mobile applications or whatever) into the workflow the administrative burden is reduced (no one had to copy numbers from one system to another) while improving customer service. It allows the plumber (small business) to focus on the plumbing and not on coordinating moving parts.

All of what I described in the example can be done now. It isn’t something for the future but something for now. The glue that binds them together are APIs. It is APIs that allow the requisite data to move between the various applications and create the seemless workflow.

The key to a digital economy is effective use of technology in small business; something that is sadly lacking in the various grand plans for digital economy from Governments and interest groups.
Reblog this post [with Zemanta]

Wednesday, May 06, 2009

RSS, Readers and Information Management

Clay ShirkyImage via Wikipedia

Several posts have looked at RSS Readers and even RSS itself questioning the need for these technologies especially with the growth in the use of Twitter and Facebook to discover news. These posts sing a premature song for the death of RSS and Readers.

RSS is a syndication protocol and is very suitable for syndicating content between applications. RSS readers are not the only use of RSS feeds. I fully expect that RSS will continue to be consumed primarily by applications rather than humans. RSS isn’t going away. It works and it is good enough for what it does.

RSS Readers are a different question. RSS Readers are primarily a way to make RSS feeds human consumable. As Dare pointed out most are based on the existing Email client paradigm. As Dare also points out this isn’t the best information paradigm for consuming large amounts of news. But this doesn’t necessitate the death of RSS Readers.

Twitter et al. are good for discovery, but not consumption. What is being seen here is a failure of filters based on time and space which no longer exist. It harks back to Clay Shirky’s comment about information overload actually being the failure of filters. What twitter et al. provide is a filter. Instead of seeing it as an either/or proposition, these filters need to be integrated into RSS Readers.

What we have is a need to evolve RSS Readers to have effective filters. RSS Readers are actually a misnomer as people focus on RSS rather than on the underlying concept of organising and presenting information. Let’s refer to the “ideal” as Information Management Application (IMA – got to love acronyms!).

IMA does the following:
  1. Gathers information (whether from RSS, Twitter, Newswires etc.)
  2. Filters and Organises information
  3. Presents the resulting information
The problem with existing methods (i.e. RSS Readers) is that they have little to no filtering and organisation. IMAs provide a rich arrangement of filters and organisation methodologies to help manage the information ocean. Like panning for gold, the IMA sieves the information ocean to find the nuggets.

It is step 2 that makes the difference. A majority of the organisation and filtering can actually be done with by grouping related sources together (e.g. rather than have 100 articles about Apple buying Twitter, group all these articles together as one), organising the stack of articles by source (e.g. the act of adding a source to the IMA makes it more important than a source not actively added to the IMA) and then layering over than filtering based on what your social network has read or is pushing.

Add in metrics about how many articles are in a group, how fast the group is growing and how much attention others are paying to the group and suddenly the ocean of news is far more manageable. Add in a touch of human curation and you have the 21st version of the personalised newspaper.

IMAs will come in multiple variants be they desktop based or online. Google has the embryonic version in Google News and Google Reader. Yahoo can also create an IMA as the next evolution of Yahoo News. The field is wide open and like many things no IMA is going to suit everyone.
Reblog this post [with Zemanta]

Tuesday, May 05, 2009

The Web & Sainsburys

Supermarket in São PauloImage via Wikipedia

This post is part of a series on how web technologies and concepts can benefit various organisations, the particular focus being APIs and open data. In this post I will look at supermarkets via Sainsburys.

Sainsburys is the UK’s second largest supermarket chain. Sainsburys wasn’t chosen for any particular reason merely it is representative of supermarkets in general. Before proceeding it is worthwhile considering what a supermarket really is. A supermarket is an organisation that aggregates shoppers by providing the convenience of purchasing a majority of consumables from one place. Larger supermarket companies can also be considered a logistics company specialising in moving groceries from producers to shoppers.
Sainsburys provides a fairly stock standard website and online shopping portal. It is rather difficult to use in my opinion and one of the key reasons that I haven’t bothered to try and shop for groceries online.

The game changes by providing APIs.

The APIs would provide access to data such as purchasing habits, aggregated demand data (e.g. 2000 tomatoes where purchased today), current stock, current prices, how much carbon/energy used in getting an item to the store. But the APIs shouldn’t merely provide data but also provide access to functionality such as making a payment, placing an order, ordering stock and communicating with the consumer.

The APIs would form the basis of Sainsburys online offerings allowing them to be easily updated and changed to allow new services and applications to arise and fall. The APIs also provide a means by which internal teams can create internal applications without the need to extensive resources. The APIs become a means to allow innovation at the edge of the organisation.

The real power of the APIs comes from allowing 3rd parties to use the APIs. Suddenly startups and entrepreneurs can create new applications that mash Sainsbury data and functionality with other data and functionality creating new value. Sainsburys creates an ecosystem of functionality based around it. The APIs allow Sainsburys to move from being simply a supermarket chain to a shopping platform that blends bricks and mortar with the web.

Many will say that the information is valuable or would provide Sainsburys competitors with an edge. They would be wrong. Information is only valuable when it is useful. Lots of the information contained in Sainsbury only becomes useful when it is unlocked. Take for example my weekly grocery list, while that is locked in Sainsburys it isn’t useful. It has no value, but provide that information to me along with items I can substitute to reduce my grocery bill then there is value, which leads to the second requirement to be useful, which is context. Information is only useful in context with other information.

The competitive edge one is very traditional thinking and it ignores the reality that the competitive advantage accrues to the company that is more open rather than less. The ecosystem allows new resources to be devoted to creating new applications, far more than Sainsburys could muster on their own. The ecosystem creates a positive feedback loop that is hard to disrupt creating advantage. Rather than Sainsburys having to devote resources to picking winners for new applications, the ecosystem does that for them. Those that create value for the ecosystem survives while those that don’t wither away.

Here is a short list of potential applications that become possible with the APIs and open data:
  • Review my shopping list
  • Subscribe to regular grocery delivery
  • Import my shopping list & prices into another app
  • Local supplier Dashboard
  • Tracking energy and carbon of purchases
  • Mobile application for navigating stores
  • Mobile application for store management
Let’s consider a few of these in more detail.

Review My Shopping List
The user gets access to all their purchases and the prices they paid. This is then reviewed looking at what could be substituted to bring the price of the weekly shop down. This could be something done by Sainsburys or could be done by 3rd parties. Probably both.

The review My Shopping would more than likely be part of a broader finance and shopping management application, but even on its own would be valuable to many people. Extensions of this basic application is to allow users to see suggested recipes based on what they purchase, start with a set of recipes and build a grocery list and set a budget and build a grocery list based on the budget and standard needs.

Local Supplier Dashboard
The dashboard would in effect allow local suppliers to provide Sainsbury with produce to particular stores using a JIT framework. The dashboard would allow the supplier to enter what they have in stock and then they can bid on meeting the stores needs for the next day.

Suppliers would also get access to information on patterns in demand allowing them to change their own production to match demand better.

Mobile Application for Navigating Stores
Each store is different and their stock availability, to say nothing of location changes all the time. The application would use the APIs to provide accurate and up-to-date location and availability for items within a store and provide directions to the location.

The application can also serve to provide a feedback loop for the store, as it would allow the user to enter information about the quality and similar information for the items.

In The End
The APIs open up a lot of opportunities and create an ecosystem based around Sainsburys. This ecosystem becomes an advantage that is very hard to disrupt. By creating value for their customers and suppliers Sainsburys will be creating value for themselves. Value that is sustainable.

Reblog this post [with Zemanta]

Wednesday, December 10, 2008

Turning Book Publishing Inside Out

A printing press in Kabul, Afghanistan.Image via WikipediaContinuing the series of collaboration and coordination platforms (C&C platforms) I wanted to look at a more concrete example of how cc platforms can change an industry and how they work to turn existing businesses inside out. My example will look at book publishing.

Book publishing has a range of steps to produce an end result (a book in the shops that is purchased by a consumer). Let’s look at the steps:

  1. Author a book
  2. Find a publisher willing to publish your book
  3. Edit the book
  4. Make changes to your book
  5. Create Cover art
  6. Print and bind book
  7. Market Book
  8. Ship book to stores
  9. Distribute revenue

Now I realise it isn’t necessarily as smooth as the list makes out and some of the items happen in parallel, but for purposes of the example it works. Steps 3 to 9 are currently the realm of the publisher. That arose as a publisher was the only one who could organise the resources needed to complete steps 3 to 9 at reasonable cost.

But with Collaboration and Coordination platforms this is no longer true. C&C platforms reduce the costs of coordinating the various activities that you don’t need a vertically integrated publisher to achieve a sold book.

For an author to publish their book, the C&C platform would provide them with access to a market of editors, designers to create cover art, print-on-demand services, book marketing services and payment features. The C&C platform doesn’t provide all the services rather it organises them and reduces the transaction costs of using services from multiple providers.

Say I am an author of a book. I upload my manuscript to the C&C platform. I then begin the workflow of publishing the book with the system creating an alert to editors that a new manuscript is ready for editing. The editor that I select then makes the edits and uploads the changes to the manuscript; I review and accept or push-back on until we arrive at something both are happy with. I then need to create the cover art and the system again creates an alert for designers. I select a designer and they get access to the manuscript so they can create relevant cover art. Once the cover art is agreed, the art is uploaded to the system and I progress to the next stage.

I now have a book ready to go, so I need to pass the manuscript and cover art along to a POD provider. This I do with a simple click on a button loading the book into the system of the POD provider I have selected. This also puts advertising and listing of the book into key retailers and I can begin the marketing of the book. It may be that I need support for the marketing which I can source through the platform as well.

The C&C platform guides the user through the steps necessary to complete a task (publishing a book), handles the necessary communications, makes sure everyone is paid and manages the media. It coordinates the various markets for each service so that together they can achieve a task. C&C platforms reduce the transaction costs to the point that a Firm in the Coaseian sense is not the most effective manner in achieving a task.

Reblog this post [with Zemanta]

Thursday, November 13, 2008

In support of James’ Cloud

Clouds rear to crashImage by Simon Cast via FlickrJames Governor recently did a re-run of his 15 Ways to Tell Its Not Cloud Computing with a post addressing some of the critical reaction he has got in a post 15 Ways I Am Wrong About Enterprise Cloud Computing. I broadly support James’ original thesis and I’ll explain why.

Definitions of Cloud Computing abound and some are really wordy and well confusing. When I think of Cloud Computing I keep it to the following definition:

“Cloud Computing abstracts the where and how of computing to allow users to focus on the what”


By this I mean that developers no longer need to worry about the details of how computing is delivered or where the computing is located (i.e. what server) instead they can focus on making sure their application achieves what they want and is reliable.

So a Cloud is more than simply a grid or utility computing as it also needs to support software stacks without the developer worrying about how it is done. A full on cloud negates many of the low level management requirements and simply provides computing and storage resource that is on-demand and easy to use, like booting an OS and running an application.

Now what we will have is internal and external clouds to an enterprise. Think Internet versus Intranet. The reason for deploying an internal cloud is to reduce the hardware capital inefficiency most enterprises face along with providing internal developers access to the benefits of Cloud Computing will still meeting the desire for security and control of data and applications.

Companies like Sun, IBM and HP will rollout “cloud-in-a-box” that allow Enterprises to replace existing hardware with Clouds. Will the enterprise own each server that makes up the “cloud-in-a-box”? Probably not. Instead they will own the “cloud” with a maintenance contract that sees the vendor swap out the hardware regularly to keep the computing capability of the cloud growing.

The reason for enterprises to deploy internal clouds is simple – it increases capital efficiency of IT while allowing developers and system administrators to focus on the application and less on keeping a mass of hardware and low-level software running and up-to-date. External clouds will work for many businesses that don’t need massive internal applications. It isn’t a really an either/or proposition.

James notes at the end of his reply they he is half-right and half wrong. I agree with the caveat that I think he is more than half-right and less than half-wrong. Until enterprises can tick off his 15 points they will not be taking full advantage of the potential of Cloud Computing.

Reblog this post [with Zemanta]

Monday, November 10, 2008

Optimisation of Workflow and Collaboration Platforms

View of Vale of Blekeley from Uley BuryImage by Simon Cast via FlickrOptimisation of workflow is the aim of the game. We want to reach the goal with as little expenditure of resources as possible. Unfortunately, much of the optimisation game has been played at the level of the individual action, which usually results in a destabilised system. This post in the series on collaboration and workflow will look at how workflow should be optimised and the role that collaboration platforms can play.

Reasonable question to start with is why optimise? Optimising or improving workflow increases the throughput. More simply, optimising workflow means more gets down with fewer resources. For business it means they can focus on producing the most value without wasting resources.

Optimisation of workflow is not about making a single action overly productive but instead about balancing the various actions in a workflow to produce the best overall throughput. It is about making the system robust rather than optimising for a particular scenario. “The Goal” by Goldratt provides a useful case study on optimising workflow.

Actions within a workflow can be re-arranged, removed, melded together and improved. The key is the modification of actions within the workflow all needs to focus on improving the overall workflow throughput. This may mean that while an individual action’s throughput can be increased from say 80 to 90%, there is no point in doing this if it does not increase the overall throughput of the workflow.
So where do collaboration platforms come in? Collaboration platforms have two functions (1) they serve as a framework within which to improve workflow and (2) they offer a way of improving individual actions.
The improvement of individual actions is a tried and tested use of collaboration frameworks. Think parallel editing of a client document and the management of tasks for a project. The improvement of individual actions is a well developed use of collaboration platforms but this only works so far as it does not optimise the action in context of the wider throughput of the workflow in question.

The framework aspect of collaboration platforms is very under developed and in terms of overall impact on business this is where changes will have the most dramatic impact on a workflow. The collaboration framework allows users to optimise and control actions within the context of the overall workflow.
The idea of a framework is to allow users to build a workflow from individual actions, examine how this workflow works and selectively change (add, remove, optimise) actions all with the aim to improve the overall workflow. It is like using plant control software to change various flow rates and values in order to change the amount of a chemical produced.

To illustrate what I mean let us look at the example of putting on an event. An event requires the coordination and completion of a series of actions such as booking and managing the venue, managing attendance and paying various entities. Each of the activities would be arranged as required into a workflow with the collaboration platform ensuring the smooth handoff between activities. None of the activities need to be powered by the platform rather they are coordinated and controlled using the collaboration platform in order to achieve the goal of the workflow. Using real world companies the venue would be booked and managed through BookingBug, RegOnline handles the registration and attendance management, Moo.com produces the tickets and ID, PayPal is used to manage payments to various entities and Huddle coordinates all these activities and manages the communication and information between the event organisers. The event organiser can focus on creating a compelling event that runs smoothly.

It is the coordination and control of actions that produces the dramatic improvement in workflow and consequently the value created by the business. As workflow now and increasingly extends across multiple organisations coordination is key to ensuring that workflow is as effective as possible. At the same time, as collaboration platforms improve the coordination of workflow then increasingly workflow will become made up of various groups working on the actions to which they add the greatest value (think of the example above). It is a positive feedback back cycle – improved collaboration increases the value of various groups working together which in turn drives improvement in collaboration and so on.

We are seeing the rise of workflow specific platforms such as Amiando and RegOnline in the case of event management. I suspect this trend will continue but there is a lack of flexibility to this approach. The real revolution will happen as the current crop of collaboration platforms along with new entrants evolve towards workflow coordination platforms that support the plug-in and specialist modules. Most groups require more than a single workflow to operate and I expect single workflow services such as Amiando and RegOnline will work within generalised coordination and collaboration platforms.

Reblog this post [with Zemanta]

Understanding Workflow

In the first post, Collaboration Reformation, of this series on collaboration and workflow, I made the statement about the transformation of collaboration from resources to activities. It is worthwhile looking at activities, or more accurately workflow, specifically before moving on to further discussion on collaboration and workflow.

Workflow is essentially a series of discrete actions that when group together produce a desired outcome. The obvious example is a manufacturing assembly line. A series of actions such as screwing on a door and adding an engine are arranged together in a line in order to assemble a car. Manufacturing assembly lines are obvious but don’t become hooked on the assembly line example. Workflow is simply a series of discrete actions performed together to achieve a goal. The goal can easily be the development and roll out of new features for a web service as it is the assembly of a car.

Goals can range from a product or service (say a car or a massage) to something more intangible (say increasing support for a candidate). For most of the post I will focus on the product and service goals but it applies equally to intangibles as well.

Workflow can be broken down into three categories: operational, development and overhead. Operational is the workflow that delivers the goal. Development is workflow that is necessary to create, improve or fix a goal. Overhead is workflow necessary to keep the organisation and group going in order that it can achieve the operational and development workflow. There is overlap between the three categories of workflow and you can represent it as a Venn diagram.

The workflows necessary to complete a goal are unlikely to be fully contained within a single group. A car maker doesn’t make the bolts or wires that go into a car. Toyota recognised this which is why The Toyota Production System works to coordinate workflow in suppliers not only within a Toyota plant.

Each workflow is made up of lots of different actions. Optimising an action without consideration for the workflow de-stabilises the whole workflow and produces counter-productive results. It is no good having one action of a workflow produce more than can be processed by downstream actions. In “The Goal” Goldratt provides a series of good examples of what happens when optimising a single action versus optimisation across the workflow.

In the next of this series, we’ll look at how optimisation can be achieved and look in more detail at how collaboration platforms are a part of this optimisation.

Reblog this post [with Zemanta]

Sunday, November 09, 2008

Keywords from Questions

In a recent article Google Search Quality Tech Lead Daniel Russell talks about an example of a user using keywords to find ferry timetable. What struck me as interesting was how the user didn’t hit upon using the keyword “ferry” until further in through their search task.

I suspect this was caused by the user starting with a question with words to the effect of “When does the ferry leave San Francisco to Larkspur?” and then attempting to turn this into a series of keywords by knocking out works such as does, when etc. The word “ferry” got knocked out of the user’s first run of keywords as it was a generic reference to ferry. In this case ferry was thought of as a common noun rather a proper noun.

If my hypothesis is reasonable then quality of keywords is going to depend on how the user first structures the question in their mind. For example if the user had used the following structure for the question “When does the San Francisco to Larkspur Ferry leave?” the word “ferry” would have been used as a keyword.

The potential importance to the way a user structures their initial question mentally points to a severe limitation to keyword and ranking search paradigm. The speed and quality of the search experience is heavily dependent on the user structuring the initial question so as to readily identify effective keywords, something that the search engines can do little to effect.

On the other hand question and fact paradigm based search engines, such as True Knowledge, will not suffer this problem.

Reblog this post [with Zemanta]

Friday, November 07, 2008

Micro-startups in a Collaboration and Coordinating world

SV200709Image by Simon Cast via FlickrIn Jason Calacanis’s recent email missive, he explores where the value for start-ups lie. Two points “The Age of the Micro-startup” and “The Try Everything Era” touches on the long on-going battle between features versus products. Put succinctly, many start-ups are little more than a feature (albeit useful) and in and of itself not a sustainable product.

Jason’s theme is that the high capital efficiency of today’s web will allow features to blossom and expand on existing services. I remain sceptical of feature companies’ (micro-startup in Jason’s terminology) as a standalone going concern but I see the value that these micro-startups can create in a collaboration and coordination world.

Collaboration and coordination platforms will enable micro-startup’s to create value by increasing the overall value of the platform by adding functionality. The advantage for the micro-startups is that the value of their feature is increased by being coordinated with various other features allowing users to achieve a goal. The value of the whole is more than the sum of its parts. Further they get access to a framework that simplifies coordination and business operations (e.g. getting revenue).

For collaboration and coordination platforms also benefit from the multiplicative effective as well as allowing features to be added to the platform cheaply and in response to demand. While they could build a lot of the features, the development resources needed that would limit what and how quickly new features can be rolled out. Micro-startups form a eco-system that is self organising about what to build and the resources to devote to it.

This differs from existing platforms, Facebook and OpenSocial, in that these platforms are about building micro-apps that have little to no coordination with other applications on the platform. The applications are standalone. These platforms essentially act as a hosting service with access to a social graph.
Coordination and Collaboration platforms are about coordinating actions (features) in order to achieve something. Achieving a goal is going to produce greater value in the long run and produce more value producing companies than simply tapping a social graph.

Reblog this post [with Zemanta]

Wednesday, November 05, 2008

Resources versus Answers – Asking a Question of Search

{{fr}} La tour Eiffel vue depuis le Champ-de-Mars.Image via WikipediaSearch is very broad in meaning and it is easy to lose sight that search actually consists of two distinct sub-sets of queries. Both sub-sets aim to find something; one is looking for resource and another for an answer. At this time we use the same approach – keywords matched in a document that is ranked for relevancy via some method (human and/or algorithm) – across both of these sub-sets of queries. This works somewhat but we are rapidly approaching the limit of effectiveness for this approach. This limit is Marissa Mayers 80/20 problem of search.

The first sub-set is finding resources (e.g. documents). The current keyword and ranking method works well for this type of query. This is what has fuelled Google’s growth. Keyword and ranking when a user is looking for one or more resources on a topic such as blog posts talking about an election. Where it falls down is answering specific queries such as “How old is the Eiffel Tower?” The user in this case is looking for a fact. Users have gotten around this problem by using the returned resources from a search as the basis to find the answer they are looking for, a human adaptation to a systemic problem.

Finding answers is the second sub-set. While we currently rely on keywords and ranking to navigate to an answer it is cumbersome and not effective. Instead the paradigm of keywords and ranking needs to be tossed out. Finding answers works better with a question and fact. A question (as opposed to queries) allows the system to quantify what fact is being asked about. For example the question “How old is the Eiffel Tower?” focuses the particular answer to be found to the age instead of potentially the location, who built it, what it is made of etc.

Using the question and fact paradigm to find answers creates new approaches to using web services and usefulness of the web to everyday life. This isn’t to say that question and fact will replace keyword and ranking rather it is complimentary and produces better results for a sub-set of search.

Consider the example of finding flights for a holiday. Using keyword and ranking the user would type in something along the lines of “flights cheap [destination]”. The engine would then return a series of web sites that match those keywords. The user then navigates to those pages and then drills through the pages to find the answer to their question. If, however, question and fact is used the user would type in “What is the cheapest flight to [destination] leaving on the 21st of December?” The web then returns the fact that flight y priced at x leaving at 10 am on the 21st is the cheapest flight. How much quicker and easier is that to understand?

For many people the web and search are still too difficult to use. But they know how to ask a question and this opens up the utility of search and the web to a whole range of users that are intimidated by it. It is worth repeating that question and fact will not replace keyword and ranking. There are queries with which question and fact doesn’t work for just as there are queries for which keyword and ranking doesn’t work for. They are complimentary.

Question and fact does have the potential to boost the growth of paid search results. The boost arises from question and fact providing a better signal of the user’s intention and so improves the targeting of advertising that better answers the query. For example a user asks the question “what is the cheapest holiday for a 16 year old girl in Mexico?” it a very reasonable to assume the intent is to find a holiday for a 16 year old girl in Mexico. A keyword and ranking would produce results about holiday’s in Mexico without any knowledge of whom or why he is searching although an assumption could be made that the person is looking for themselves. Interestingly, through in demographics and/or behavioural data and the system will produce completely the wrong answer. Say for example the person is 52 year old male in which case the system is likely to return Mexico holidays for a 52 year old man when his intention was to find a holiday for his 16 year old daughter.

Question and fact will go a long way to addressing the 20% of search remaining. Many web services implement crude methods for asking a question, ones that are frankly laborious and time consuming to use. The key to unlocking the power of question and fact is to make it as easy as possible to ask the questions. The pitfall to implementing question and fact is knowing when to use it. Question and fact works when the question can be answered by a fact e.g. “How old is the Eiffel Tower?” It doesn’t work when the answer is not a fact e.g. “What is the best holiday in Mexico?”

Reblog this post [with Zemanta]

Friday, October 31, 2008

The Collaborative Reformation

Collaboration has become the buzz word of the times. It holds out the promise of making all work productive all of the time, an attractive carrot in times of economic crisis. While collaboration tools will improve the effectiveness of existing processes and businesses its true impact and most dramatic effect lies in how collaboration tools can bring about a reformation of business.

Improving the efficiency of existing process works up to a point and indeed most, if not all, collaboration tools are predicated on somehow improving the efficiency of existing processes. The real promise of collaboration tools is how they can help users reform the fundamental processes of businesses. I’m not simply talking about getting rid of layers of approval but the complete overhaul of how new work is brought in, how it is created, how it is charged and how it is produced.

The impact comes from allowing business and users to focus on the work that adds value and streamline and eliminate the non-value add work. Elimination may involve out-sourcing to another business where the particular process or work is their value-add. Think designing cover art for a new book. It is not a value add proposition for a publisher but is a value-add proposition for a designer.

The reformation extends beyond simply managing documents and information to completing the core tasks of the business whether it is a plumber, development agency or a manufacturer. Collaboration services need to be given access to the physical world that plays such an important part in many businesses whether it is tasking plumbers and ordering plumbing supplies for delivery or controlling a CNC machine.

In effect collaboration services need to evolve into a framework within which business operates; a framework which supports agile business processes, modules for specialist features (think CAM control) and management of information within the business. This is the path for development of collaboration services such as Huddle.

The ultimate goal is supporting the ideal of the networked business. A “Business” is a network of smaller businesses using a common collaboration platform with specialist modules from various providers. Business becomes a network of networks in which the collaboration service coordinates activities. It is this that is the root of the dramatic and sustainable change that collaboration services can bring to business.

Reblog this post [with Zemanta]

Tuesday, October 21, 2008

Can Hubdub Survive?

I’ve been playing with Hubdub recently and all I can say is it has a looong way to go. In fact I think that unless some major changes are made Hubdub won’t survive 2009. Unfortunately, I think Hubdub faces some major, major hurdles, that will make the struggle to build a decent revenue stream (remember now cash is king...angel funding will only get them so far).

Hubdub for those that have not come across it is as predication market based around news. The idea is to combine news aggregation of some sort with predictive markets.

I’ve put down bets and created questions. My consequent experience has been less than heart warming. But to illustrate my concerns let’s look at the questions I created.

My first question was “Will Digg buy Hubdub in 2009?” Speculative yes, but Hubdub is a predictive market – questions are by nature speculative. Background to the question: Kevin Rose discussed Digg’s international expansion plans in his talk at FOWA London. This talk was widely reported in the media. Some other facts:

  • Digg has just closed a funding round of $29m in September
  • Hubdub has only raised angel funding and has 4 employees
  • In this economic climate cash is king. Getting revenue positive is the holy grail
  • Hubdub has not articulated a source of revenue that is sustaining
From these data points (which are all easily available with a quick search on the web) one concludes that Hubdub is a good target for acquisition. It has a decent (although it requires some work) prediction platform but other than that it has nothing special. Kevin Rose wants to expand internationally and the prediction market technology would work well with the Digg platform. It would certainly give Digg greater number of potential revenue streams. I will be the first to admit that this is speculative but it is based on facts. But all it took was one person to raise a question and the question was voided.

Second question was “How far will UK house prices fall by December 2008?” This question was voided as it didn’t have an option for housing prices fall being below 15%. Let’s look at the logic of this. As of September 2008 both Halifax and Nationwide have reported house price falls of 12.4%. There are three more months to go before the end of December and these things don’t turn around on a dime. Being below 15% is not an option as credit is still tight and the UK has entered recession. There is no sound possibility that house prices will not fall less 15%. Now let’s look at this from angle of question creator. I may not want to provide an option so why should I? Why do questions have to be modified to provide gamblers with an option they want?

The issues with the questions are merely symptoms of what I see as major flaws of the Hubdub system. They are:
  1. The rules for creating questions are vague and easily open to interpretation.
  2. The very act of creating questions is daunting and annoying
  3. The site is rapidly becoming dominated by power users
Hubdub is seeing the play out of Clay Shirky’s maxim “A Group Is Its Own Worst Enemy”. Power users and early adopters will band together to void questions and in other ways hassle newbies merely because they don’t like the question (rather than it’s predictive quality) or because newbies are falling afoul of capricious unwritten rules. There is no penalty against power users for this type of behaviour. This is over and above the effort needed to create questions in the first place. It is extremely disconcerting to put the effort into questions only to have them void on relatively spurious grounds.

The single most worrying aspect about the flaws in Hubdub is everyone has been talking about the development of communities and their interaction for the last 11 years. Let’s recap – Usenet went through the same problem, Slashdot went through the same problem, Digg went through the same problem. See a pattern here?

Clay Shirky has been shouting from the roof tops about it for years. Hugh McLeod and Tara Hunt have all discussed it. At what point do people pay attention? Angel investing or not, for a service that is based on community and users, to not have the necessary tools needed to manage the community’s interaction with the platform is simply, well, scary. It speaks to a company that has a fundamental lack of understand of community base services.

Is it important? Very. Hubdub needs a diverse, large and vibrant community not only of speculators but also question creators. The way Hubdub is going to make money is from selling premium access to data and audience to businesses. However, companies will only pay if the Hubdub community is diverse and large as then the data and audience has value to them. As it is the current community is doing very well at driving new members away. Hardly a good method of growth.

Hubdub can possibly turn this around. The first step is to build the tools and features necessary to manage the community interaction. This will piss off the power users and early adopters as it blunts their power. That is the price to pay for improving the experience and engagement for a broader and more diverse people. Actions must have consequence. When actions have no consequence poor behaviour soon dominates. Slashdot found this out the hard way as has Digg.

The other part of the engagement issue is question creation. Relying on people to read FAQs about question creation is very, well, RTFM. Most people don’t RTFM and nor should they. If there are rules about question creation they need to be clear and objective with no room for abuse to void questions that someone doesn’t like. Of course, if the rules are clear and objective the system should not allow questions to be created in the first place that don’t meet those rules. People should not have to RTFM.

In fact, I wonder why voiding is necessary at all – isn’t the very act of betting on a question a vote on the question's quality? Why not just use the activity on the questions as a way to surface or subsume questions? Activity is a much better method than voiding or voting. Using activity blunts the prejudices and power of any single person or group of people.

The interleaving idea through the emails from Hubdub and site of not being a speculative market still stumps me. It’s a predictive market they are, by definition, speculative. It’s like being a fish and trying not to drink the water. Hubdub is a speculative market – if the problem is with questions that don’t have any news or are a long time in the future have a special section for these types of questions. Don’t ban them.

I did hope that Hubdub would be good. But I am sorely disappointed and I now doubt the company’s survival. I certainly would not invest any money in the company without some major changes to the platform.
Reblog this post [with Zemanta]

Saturday, October 04, 2008

Widgets, Communities and the Edge

The web is making it easier and easier for groups and communities to form. Groups foster social cohesion by having members demonstrate affiliation and by the use of objects to create community identity. Think Star Trek fans wearing Star Trek uniforms at conventions or fans of Metallica wearing Metallica branded tee-shirts.

Unfortunately web based methods of indicating affiliation don’t really translate to the real world. This is important as groups are increasingly rooted in the real world, indeed traditional line between cyberspace and the real world is becoming increasingly blurry.

Personalisation services offer the ability to create physical objects that indicate affiliation and community identity. These services are centralised and therein lays the problem. By being centralised they impose a coordination cost on the groups.

Widgets offer services like MOO.com and Ninjazoo the opportunity to offer personalised and communitised products directly into the community without getting in the way. Widgets provide a means of removing the coordination cost on groups by meshing the service within the normal activities and sites of the group.

It is taking the mountain to Muhammad rather taking Muhammad to the mountain.

It is the distribution of core functionality where the true value of widgets lies. Not with the distribution of content but allowing web services to adjust to an Edge Economy.

Tags: MOO.com, Ninjazoo, Edge Economy, Web Services, Web 3.0

Friday, October 03, 2008

Data Half-life: Time Dependent Relevancy

Data Half-Life is not an indication of the importance of a particular piece of information. It is actually a measure of how long a piece of information is relevant. Relevance is not a substitute for importance. It is dependent on context and the information itself. So a low data half-life means that the piece of information will quickly lose its relevancy. A high data half-life means the relevancy will drop slowly.

Consider the story that Clay Shirky related in his keynote at Web 2.0 Expo in New York. In this story someone changed their relationship status from engaged to single. This information is highly relevant to some people and not very relevant to most others. Given that data half-life reflects the broader relevance of the information to a person’s network, it has a low data half-life. It is generally not relevant to most of the people in the network.

Now they many want to know or feel the need to know, that does not mean it is relevant to them. It is easy to mistake the desired to know or the need to know as relevant. Desire to know has no bearing of the information’s data half-life.

By having a low data half-life the relationship status will only travel only so far through the person’s network, thereby avoiding the result in Clay Shirky’s story. Data Half-life is represents how time dependent the information is. The more time dependent some data is, the lower the half-life and the less time dependent the higher the half-life.

Tags: Filters

Thursday, October 02, 2008

Privacy Filters and Facebook

In my previous post I used privacy in Facebook as an example of how data filters could work. One point I glossed over was how currently Facebook, indeed all social sites, fail with social distance. Unfortunately, social distance is a necessary for privacy filters to work satisfactorily.

Facebook has one major flaw, once a person is a friend in Facebook they are treated the same as all other contacts whether the connection comes from bumping into the person at a pub or someone you grew up with. It collapses the privacy or social distance between two people. The social distance can be considered how strong the connection between two people is. Social distance provides a measure of both strong and weak ties as articulated by Mark Granovetter.

Without some measure of social distance or strength of connections, any privacy filter is going to fail. The social graph fails to represent the real world connections between people properly.

Facebook attempts to use groupings of friends to approximate social distance but this is cumbersome to use. The manual nature of setting up and categorising everyone into groups is a major barrier to use. People are lazy.

What is needed is an automated method for calculating social distance. Social distance is calculated (and this is how Mark Granovetter categorised connections) by the frequency of communications. Measuring frequency of communications is difficult for Facebook. While Facebook can measure wall posts, internal emails, poking etc., so much more of our communication occurs outside of Facebook, outside of the wall; whether through email, IMs, phone calls, SMS, twitter parties attended etc.; that the frequency of communication within the wall is not a reasonable approximation for the wider frequency of communication.

The key measure of social distance – communication – is hard to quantify as it is dispersed through many different channels. Trying to capture the frequency of communication via porting the data in is one method of dealing with the issue. The other, probably more realistic, method is to start off with some rules and use what can be easily quantified to refine the measure of connection strength overtime.

The rules would look at what is known generically about social connections. Some of rules are:

  1. Married is a strong connection
  2. The same surname is a strong connection
  3. If strong connections to friends with which you have strong connections then you probably have a strong connection
Some of these rules will dictate a very strong connection (first rule) while others will dictate varying strengths dependent on factors such as prior connections with other friends (third rule). All connections start as very weak and are refined first by application of the rules and then overtime by measures of frequency of communication.

Privacy filters all start with knowing the distance between two end points whether physical in case of centuries before or by social distance in the case of today. Until Facebook and any other social-based site has a measure of social distance privacy filters are going to be mediocre at best and more often prone to failure.

Tags: Privacy, Facebook, Filters

Friday, September 26, 2008

Failure of Filters

The title from this post is taken from the keynote that Clay Shirky delivered at the NY Web2.0 Expo in September 2008. The premise of the keynote is that the “information overload” we are facing is not a problem but a fact (one that has been around since Gutenberg and his movable type press) and what we are seeing now is the collapse of the traditional filters that mediated the information overload.

The existing filters for information were founded in the difficulty of moving information over distance. The various communications technologies of the 20th Century have steadily eroded the tyranny of distance. The web completed the destruction of distance filters by removing all concept of spatial distance for information.
Our sense of privacy is again bounded up in the hassle in moving information over distance. This physical distance is the basis for the whole concept of privacy. The closer we are to other people the less privacy we expect. We found that to be a reasonable rule of thumb as those closest to us (community, family, friends) are likely to spatially close to us. We only now need privacy safeguards because the rule of thumb no longer applies – spatial distance is meaningless for information now.

Information overload and privacy issues are a rooted in us expecting that filters based on spatial distance to continue working in a world where information has no spatial component. Any filters built with this expectation don’t work. Instead we have to create a new framework for filters that don’t rely on spatial distance.

By borrowing ideas from science we can create a framework that doesn’t rely on spatial distance. The framework is based on data half-life, data permeability and data potential. Data half-life is the measure of how long the bit of data takes to lose half of its relevancy/ importance. Data permeability is a measure of how hard it is for data to move over a period of time – think fluid moving through a filter. Data potential is the initial potential for the data to move – think potential energy in Newtonian dynamics.

The interaction of these three parameters determines how far and how quickly information can travel within an environment where spatial distance has no meaning. An analogy will help illustrate how the parameters behave together to filter information.

Let’s say we have some information – death of the chief of a village. The village has good roads and the news is to be sent by horse. This information will go far as it important (chief of a village), it is easy for the information to move on the road and the horse is quick. If, however, the death is not the chief then the news won’t travel as far it is not as important. It is the interaction between the data half-life (how important the person is), the data permeability (how easy it is to move the information) and data potential (how fast the information can move) which determines how far the information will travel.

Changing the parameters creates a varied set of filters that determines how far and how fast information will defuse. Each connection has a level of data permeability with information coming in assigned a data half-life and a data potential. The information only passes the filter when the data half-life and data potential are enough to overcome the data permeability.

To illustrate consider changing your relationship status in Facebook. If someone changes their status from relationship to single they don’t necessarily want the information to spread quickly through their “facebook friends” as their friends will include work colleagues and friends of friends only met once. Instead each of their connections should have different data permeability and depending on information (data half-life and data potential) it will show up in some of the connections news feeds right away, some in days, some in weeks and others never at all.

There is no single way to create and calculate data half-life, data potential and data permeability. Various developers will come up with their own methods. Some of which will work and others that won’t. Hopefully further down the track we will see a standardisation on calculating the parameters based on accepted criteria for each type of information – personal, communications, knowledge etc.

Tags: Filters, Information Overload, Privacy, Clay Shirky

Sunday, January 06, 2008

Language and problem solving

In the most recent New Scientist (Vol 197 No 2637) there is an interesting article discussing the issue of language and how it frames problems. The perspective of the article was that English's newtonian way of describing the world failed to frame questions properly for quantum and other similar non-newtonian physics. The article even goes so far to say that the lack of progress in non-newtonian physics is because problems are framed via the language with a newtonian world view.

Does the same problem exist in the world of the internet? While I realise the Internet world is great at creating new words, these are still framed by the overall language. A language that is "newtonian". As Internet shifts to flows and systems as opposed to objects and links, do we need to look at how we frame the discussion via language to open up the problem solving juices of the internet community? New next wave of innovation will be less around nouns towards verbs, the doing rather than the being and yet we still primarily use nouns in discussing the web and its evolution. Should verbs that describe process, systems and flow be the primary descriptors of the next web?

The article describes an example of Montagnais phrase "Hipiskapigoka iagusit". It very, very roughly translates to "singing health", a process, within which a medicine man and sick person exist. However, a dictionary written in 1729 translated into something that emphasised the objects and not the process. The web is shifting to loosely coupled processes as opposed to objects. I wonder whether the discussion of Robert Scoble's recent tiff with Facebook, would have evolved differently if the language emphasised process (say maintaining contacts) as opposed to data (the contacts themselves). The discussion was about who owned what objects (the contact data) rather than what the ins and outs of maintaining contacts. Another example is the current discussion going on about whether data is a commodity or not. Again the language is of objects rather than flow. How would this discussion evolve if it was frame by a language of flow (verbs) as opposed to objects (nouns)?

The same questions can be asked of programming. Everyone expresses the need to ramp up parallel programming to take advantage of the distributed nature of the internet and multi-core processes. However, can any real problem be solve properly while the language used to frame the problem is based on objects rather than flow? Does the conceptual framework that underpins object orientated programming preclude successful problem solving in the parallel world? Yes there are languages that focus specifically on parallel programming but I am also talking about the language used to describe and communicate the problem. These will need to respond to the requirements of a parallel world for people to solve problems and communicate solutions.

A lot of questions asked. I don't have the answers and I expect no one will for a while. It is interesting to step away from objects and consider things from a flow perspective. I even think I need to re-visit my recent post of Data Ecosystems and look at it from the perspective of flow rather than objects

Tags: Data, Language, Programming, Internet, Physics, Data Ecosystems

Friday, December 14, 2007

Web Next & Data Ecosystems

Web Next is not some quantum leap in reality but rather the culmination of several long term trends. Web 1.0 was the translation of real world services (e.g. Amazon) onto the Internet and Web 2.0 is Darwinian evolution of UI and social media tools and technologies. Web Next is exploiting of information to achieve new products and services with no direct analog in the real world. Some will call this Web 3.0 but I prefer the simpler moniker Web Next.

Web Next is about the creation of value not through the control of information but via the creation of synergies and knowledge through combining information and functionality. There already exists the primitive examples in the Web2 world, those such as the map-based mash-ups. Essentially value is derived via the showing a spatial relationship between data. However, these mash-ups are relatively primitive. They rely on a users existing knowledge of the spatial area in question. I personally have no appreciation of the real layout of New York City having never been there. Consequently, the value I gain from viewing or using a mash-up consisting of crime statistics plotted on a map of NYC is less than someone who has visited which will be less than a resident of NYC.

The synergy of information and functionality is created through Data Ecosystems.

Data Ecosystems


A Data Ecosystem is two or more different data sets that when combined with complementary functionality produce multiplicative effect in usefulness. Or put another way, a Data Ecosystem contains more than one source of data (a data set) (e.g. temperatures and rainfall) that can be combined, analysed and processed with the overall Data Ecosystem being more valuable than data or functionality on its own. Importantly, having more than one data set is not sufficient on its own. Rather you need various tools and functions that allow the user to act on the data. An example will help to clarify.

Take an individual piece of data, say a series of temperatures measurements. On its own you can't do much with those temperature measurements but combined with rainfall measures, annual growth rates and a map, those temperature measurements suddenly have a lot of value. Now a farmer can research and plan when to plant his crops or adjust his crop forecasts based on historical growth rates versus temperature and rainfall. To be able to make the forecast of crop tonnage the farmer needs a series of tools that allow him to find correlation factors and extrapolate the growth trends based on rainfall and temperatures. Without those functions having the data is not particularly useful. There is no use in having gobs and gobs of data if there is poor functionality in the Data Ecosystem.

Data Ecosystems highlight a very interesting point about data. One that I find is continually ignored or not understood by most data companies (including web companies). Data on its own has little intrinsic value. Data only has value with what you can do with it and what you can do with it is determined by what other data you have along with the functionality you can apply to the data.

Like a biological ecosystem, a data ecosystem must mesh together. A Data Ecosystem needs to be internally consistent. If a Data Ecosystem is not consistent then it will not generate value for the user. There is no point in trying to create a Data Ecosystem that has rainfall patterns from Australia and crime statistics in New York City. Designers of Data Ecosystems must not design the systems so they become inconsistent. Consistency is crucial. But given human nature I fully expect consistency will be ignored.
"And there's the sign, Ridcully," said the Dean. "You have read it, I assume. You know? The sign which says 'Do not, under any circumstances, open this door'?"
"Of course I've read it," said Ridcully. "Why d'yer think I want it opened?"
"Er...why?" said the Lecturer in Recent Runes.
"To see why they wanted it shut, of course."
-Terry Pratchett, Hogfather

At this point I expect some readers will be thinking that Data Ecosystems is simply the Semantic Web. Data Ecosystems is not the Semantic Web. Semantic Web technologies will be a part of Data Ecosystems, but Semantic Web is neither a precondition for nor sufficient on its own in order to build Data Ecosystems. Semantic Web helps by automating building the relationships between bits of information. In the temperature example above Semantic Web would have provide information such as the lat/long of the measurements, how it was measured, the accuracy of the measurement, the date and times of the measurement etc. This would then allow a computer to match the data automatically with rainfall data from the same location and time and plot together on a map. Semantic Web makes building Data Ecosystems easier and like objects in programming will allow Data Ecosystem platforms to increase the ease the deployment of Data Ecosystems.

Data Ecosystems can also be built of other Data Ecosystems. The output of several Data Ecosystems can be used as the sources for another Data Ecosystem and so on. Each step creating more value by allowing an individual to achieve more. Data Ecosystems will in effect create an L-space

Why are Data Ecosystems Important?


Data Ecosystems are important for one very, very crucial reason. Data Ecosystems allow people to achieve things effectively. Unlike Web 1.0 which was essentially removing transaction costs from existing real-world processes, Data Ecosystems unlock the potential for new services that are impossible in the real world.

Consider a Data Ecosystem based travel service. Such a service will allow you to research a holiday; book all transport, accommodation and activities; create a comprehensive itinerary of the holiday, send alerts at key points along the trip; calculate how much money you'll spend on the holiday; help you automatically tag video, audio and photos from the holiday and create holiday memorabilia from the items you have uploaded. All through a single Data Ecosystem.

And there are hundreds, thousands, millions of probable Data Ecosystems that have no analog in today's web.

How Things are Already Changing


To close out I want to consider something that has been banging its way around the blog-sphere and offline world: the fate of journalism. Without re-hashing the debate you can read Bill Keller's speech with Jeff Jarvis's responses here and here as background.

If everyone is a citizen journalist, then what is the point of professional journalist? A seemingly valid question but one that has the implicit assumption that both are or will be doing the same process. From the perspective of Data Ecosystems the job of a professional journalist becomes very different from a citizen journalist. The citizen journalist is a source of data. They will most likely only provide a very narrow bit of data on any particular story. Put another way, the citizen journalist becomes a source like the news wires.

The professional journalist moves on from being the source of the story to gathering all the disparate bits of information about a story and then assembling into a consistent and cohesive context around the core story. Professional journalists go form being the gate keepers to information to value builders by creating context to stories. The role of a newspaper/media company is to provide or create a Data Ecosystem within which the professional journalist can assemble, create and publish the context to stories. Within Data Ecosystems, professional journalists, news agencies and citizen journalists will co-exist and combine to produce a more valuable service than either would on their own or exists today.

The media world is already going through pain as it is forced to adjust to the realities of an information economy. Data Ecosystems provide a means to effectively adapt to the information economy. But Data Ecosystems are not limited to media. Data Ecosystems will exist right across the information world. In fact they will reach into material world as L-space is linked into materials at the molecular level. This is the true revolution of Data Ecosystems, they facilitate the merger of L-space and Real Space.

Tags: Web Next, Data Ecosystems, Web 3.0, Web Services, Semantic Web, Web 2.0, L-Space