Showing posts with label Search. Show all posts
Showing posts with label Search. Show all posts

Wednesday, May 06, 2009

RSS, Readers and Information Management

Clay ShirkyImage via Wikipedia

Several posts have looked at RSS Readers and even RSS itself questioning the need for these technologies especially with the growth in the use of Twitter and Facebook to discover news. These posts sing a premature song for the death of RSS and Readers.

RSS is a syndication protocol and is very suitable for syndicating content between applications. RSS readers are not the only use of RSS feeds. I fully expect that RSS will continue to be consumed primarily by applications rather than humans. RSS isn’t going away. It works and it is good enough for what it does.

RSS Readers are a different question. RSS Readers are primarily a way to make RSS feeds human consumable. As Dare pointed out most are based on the existing Email client paradigm. As Dare also points out this isn’t the best information paradigm for consuming large amounts of news. But this doesn’t necessitate the death of RSS Readers.

Twitter et al. are good for discovery, but not consumption. What is being seen here is a failure of filters based on time and space which no longer exist. It harks back to Clay Shirky’s comment about information overload actually being the failure of filters. What twitter et al. provide is a filter. Instead of seeing it as an either/or proposition, these filters need to be integrated into RSS Readers.

What we have is a need to evolve RSS Readers to have effective filters. RSS Readers are actually a misnomer as people focus on RSS rather than on the underlying concept of organising and presenting information. Let’s refer to the “ideal” as Information Management Application (IMA – got to love acronyms!).

IMA does the following:
  1. Gathers information (whether from RSS, Twitter, Newswires etc.)
  2. Filters and Organises information
  3. Presents the resulting information
The problem with existing methods (i.e. RSS Readers) is that they have little to no filtering and organisation. IMAs provide a rich arrangement of filters and organisation methodologies to help manage the information ocean. Like panning for gold, the IMA sieves the information ocean to find the nuggets.

It is step 2 that makes the difference. A majority of the organisation and filtering can actually be done with by grouping related sources together (e.g. rather than have 100 articles about Apple buying Twitter, group all these articles together as one), organising the stack of articles by source (e.g. the act of adding a source to the IMA makes it more important than a source not actively added to the IMA) and then layering over than filtering based on what your social network has read or is pushing.

Add in metrics about how many articles are in a group, how fast the group is growing and how much attention others are paying to the group and suddenly the ocean of news is far more manageable. Add in a touch of human curation and you have the 21st version of the personalised newspaper.

IMAs will come in multiple variants be they desktop based or online. Google has the embryonic version in Google News and Google Reader. Yahoo can also create an IMA as the next evolution of Yahoo News. The field is wide open and like many things no IMA is going to suit everyone.
Reblog this post [with Zemanta]

Sunday, November 09, 2008

Keywords from Questions

In a recent article Google Search Quality Tech Lead Daniel Russell talks about an example of a user using keywords to find ferry timetable. What struck me as interesting was how the user didn’t hit upon using the keyword “ferry” until further in through their search task.

I suspect this was caused by the user starting with a question with words to the effect of “When does the ferry leave San Francisco to Larkspur?” and then attempting to turn this into a series of keywords by knocking out works such as does, when etc. The word “ferry” got knocked out of the user’s first run of keywords as it was a generic reference to ferry. In this case ferry was thought of as a common noun rather a proper noun.

If my hypothesis is reasonable then quality of keywords is going to depend on how the user first structures the question in their mind. For example if the user had used the following structure for the question “When does the San Francisco to Larkspur Ferry leave?” the word “ferry” would have been used as a keyword.

The potential importance to the way a user structures their initial question mentally points to a severe limitation to keyword and ranking search paradigm. The speed and quality of the search experience is heavily dependent on the user structuring the initial question so as to readily identify effective keywords, something that the search engines can do little to effect.

On the other hand question and fact paradigm based search engines, such as True Knowledge, will not suffer this problem.

Reblog this post [with Zemanta]

Wednesday, November 05, 2008

Resources versus Answers – Asking a Question of Search

{{fr}} La tour Eiffel vue depuis le Champ-de-Mars.Image via WikipediaSearch is very broad in meaning and it is easy to lose sight that search actually consists of two distinct sub-sets of queries. Both sub-sets aim to find something; one is looking for resource and another for an answer. At this time we use the same approach – keywords matched in a document that is ranked for relevancy via some method (human and/or algorithm) – across both of these sub-sets of queries. This works somewhat but we are rapidly approaching the limit of effectiveness for this approach. This limit is Marissa Mayers 80/20 problem of search.

The first sub-set is finding resources (e.g. documents). The current keyword and ranking method works well for this type of query. This is what has fuelled Google’s growth. Keyword and ranking when a user is looking for one or more resources on a topic such as blog posts talking about an election. Where it falls down is answering specific queries such as “How old is the Eiffel Tower?” The user in this case is looking for a fact. Users have gotten around this problem by using the returned resources from a search as the basis to find the answer they are looking for, a human adaptation to a systemic problem.

Finding answers is the second sub-set. While we currently rely on keywords and ranking to navigate to an answer it is cumbersome and not effective. Instead the paradigm of keywords and ranking needs to be tossed out. Finding answers works better with a question and fact. A question (as opposed to queries) allows the system to quantify what fact is being asked about. For example the question “How old is the Eiffel Tower?” focuses the particular answer to be found to the age instead of potentially the location, who built it, what it is made of etc.

Using the question and fact paradigm to find answers creates new approaches to using web services and usefulness of the web to everyday life. This isn’t to say that question and fact will replace keyword and ranking rather it is complimentary and produces better results for a sub-set of search.

Consider the example of finding flights for a holiday. Using keyword and ranking the user would type in something along the lines of “flights cheap [destination]”. The engine would then return a series of web sites that match those keywords. The user then navigates to those pages and then drills through the pages to find the answer to their question. If, however, question and fact is used the user would type in “What is the cheapest flight to [destination] leaving on the 21st of December?” The web then returns the fact that flight y priced at x leaving at 10 am on the 21st is the cheapest flight. How much quicker and easier is that to understand?

For many people the web and search are still too difficult to use. But they know how to ask a question and this opens up the utility of search and the web to a whole range of users that are intimidated by it. It is worth repeating that question and fact will not replace keyword and ranking. There are queries with which question and fact doesn’t work for just as there are queries for which keyword and ranking doesn’t work for. They are complimentary.

Question and fact does have the potential to boost the growth of paid search results. The boost arises from question and fact providing a better signal of the user’s intention and so improves the targeting of advertising that better answers the query. For example a user asks the question “what is the cheapest holiday for a 16 year old girl in Mexico?” it a very reasonable to assume the intent is to find a holiday for a 16 year old girl in Mexico. A keyword and ranking would produce results about holiday’s in Mexico without any knowledge of whom or why he is searching although an assumption could be made that the person is looking for themselves. Interestingly, through in demographics and/or behavioural data and the system will produce completely the wrong answer. Say for example the person is 52 year old male in which case the system is likely to return Mexico holidays for a 52 year old man when his intention was to find a holiday for his 16 year old daughter.

Question and fact will go a long way to addressing the 20% of search remaining. Many web services implement crude methods for asking a question, ones that are frankly laborious and time consuming to use. The key to unlocking the power of question and fact is to make it as easy as possible to ask the questions. The pitfall to implementing question and fact is knowing when to use it. Question and fact works when the question can be answered by a fact e.g. “How old is the Eiffel Tower?” It doesn’t work when the answer is not a fact e.g. “What is the best holiday in Mexico?”

Reblog this post [with Zemanta]

Friday, December 21, 2007

How not to force people to unsubscribe (This means you Spock)

Today I received an email from Spock for someone trying to add the company recruitment email to his "Web of Trust". Fine people harvest email address all the time. What annoyed me is that it required two entries of the email address and clicking on a link in an email.

Not good. Made the job of unsubscribing time consuming. And it wasn't like I was a registered and the email was from someone who had found me but rather spam contacts email. Long story short, do not use Spock. If this what the company puts you through just to unsubscribe from contact spam then I hate to think how annoying the service would be if you are a registered user.

Point to all companies. Unsubsribe must be a single action on the part of the user. Double, triple, quadruple actions are a no-no. And it musn't take 10 days to filter through your system. That is a load bullshit. All it means is that you can't be arsed to fix you email marketing system and you are going to try to spam me as much as possible in that 10 days.

Tags: Spock, Email Marketing,, Email

Thursday, April 05, 2007

For search startups the real barrier to entry is the index

I was drawing up a brief document for my CEO yesterday around search, the developing technologies and the startups in the space. What caught my attention was the number of starts calling themselves "Google-Killers". It took me a while to put my finger on it but I just didn't agree no matter how cool/ground breaking/esoteric their technology was.

Why don't I see these companies such as Powerset or Hakia as "google-killers"? Index. The simple fact is that unless the company has an index that is significant percentage of the Google or Yahoo or Microsoft search index's they can't compete. Early on the size of the web was such that a new search engine could easily develop a useable index. Now it is orders of magnitude harder to not only develop the index but also make it fast, reliable and generally useful.

There is a way to mitigate this barrier to entry and that is to pick a vertical and index that. Indexing a vertical is a much easier that trying to index usable portion of the internet. Given the nature of the technologies that Powerset, Hakia et al are deploying I think health (which is also relatively open and new) would be good fit for them.

[I had included Yedda as a Google-Killer but as Yaniv Golan pointed out in the comments, Yedda is along the lines of Yahoo Answers with knowledge ranking.]

Tags: , , , , ,