May 20, 2007

Branding ads and ad listings

I just saw this post on BoingBoing about this ars technica post on the psychology of banner ads which is interesting in light of Google's and Microsoft's recent acquisition of 'creative/banner ad' networks.

"The research concludes that repeated exposure to a product via banner ads generates a positive feeling towards that product. The good news for consumers is that a critical reevaluation of the product can make these positive feelings vanish." and
"This suggests that familiarity-based advertising may work best for impulse buys, where more detailed evaluations aren't likely to occur."

I found these two papers yesterday when reading up on economic theory and internet advertising. This one Internet Advertising and the Generalized Second-Price Auction - Selling Billions of Dollars Worth of Keywords is the most readable of the two and talks about the bidding process for online ad placement. They cover the history and describe the shift from "pay what you bid" to the current "pay the next lower bid" (called a 'generalized second price auction'). Really very interesting stuff.

The second paper Brand and Price Advertising in Online Markets looked at different fundamental forms of advertisements - brand advertising and 'price' advertising. I had high hopes for learning from it, but it was very dense with too much lingo specific to this research area for me to fully understand. Here's an example: In contrast to models where loyalty is exogenous, these crosschannel effects lead to a continuum of symmetric equilibria. Yeah, I, uh, was thinking the same thing.

Their models and assumptions also seem questionable and so I don't know that their results apply to the world we live in, as opposed to the simplified models they used to prove their theory. In any case, the questions being asked and the attempt to find answers are still valuable.

Here are their findings (and again, these may not really apply to the real world)
While each firm finds it optimal to advertise its brand in an attempt to “grow” its base of loyal customers, in equilibrium, branding (1) reduces firm profits, (2) increases prices paid by loyals and shoppers, and (3) adversely affects gatekeepers operating price comparison sites. Branding also tightens the range of prices and reduces the value of the price information provided by a comparison site.
Their research shows that brand advertising allows a firm to have higher prices, since loyal consumers aren't as price sensitive, but their conclusions are that profits are less - which I don't understand, unless the cost of creating the brand is very high. The Ars Technica post shows that other research continues to confirm that viewing brand ads creates a positive impression, which is one step towards converting shoppers into loyal customers.

This got me thinking about the various forms of advertising. The second paper distinguishes 'brand advertising' from 'informational advertising', which I agree is a useful distinction. I think of it in terms of how actionable the advertising is - how delayed is the payoff. For branding, the payoff is very indirect but could be profitable (or not, depending on which research you subscribe to) due to higher prices or repeat business or cutting out the competition through causing the customer to not engage in comparison shopping. For ad listings that show a very specific product or category, which are often a gateway to a purchasing decision, the payoff is fairly direct. Some even feel that the advertising cost model may evolve into 'cost per action' and go beyond 'cost per click'. However, cost per click is currently much easier to gather metrics in a two-way trusted fashion than measuring the final transaction in a two-way trusted fashion.

Another example of a very actionable advertisement is the Amazon or EBay offer listings - these are immediately purchasable, and through syndication via Associates are widely used as advertisements. I don't know of any company that does placement optimization of Amazon Associate links (raising placement for offer with better click through rates, better commission for the associate, etc), but I think Amazon has started doing some of that with Omakase links. One interesting thing to consider is that both Amazon and EBay are similar to price comparison sites due to the large number of offers for a single authoritative item (EBay doesn't have very good item authority, but people work around that issue by using manual searches). Getting top placement on an Amazon offer listing page, or in the 'buy box' on the details page, doesn't use an auction bidding process the way Google or Yahoo paid search listings do. The offer listing position is based on the offering price and estimated shipping costs.

May 19, 2007

Online Ad industry consolidation

So, what's up with all the acquisitions of online ad networks within the past 30 days?
  • Google purchased DoubleClick for $3.1B on 4/13/2007 (on revenues of $150M).
  • Yahoo purchased the remaining interest in RightMedia for $680M on 4/29/2007 (on revenues of $70M) They had purchased a 20% stake back in Oct 2006.
  • Microsoft acquired European mobile ad network ScreenTonic on 5/3/2007.
  • Microsoft acquired aQuantive for $6.1B on 5/18/2007 (on revenue of $442M).
  • AOL acquires major interest in Adtech AG on 5/16/2007.
  • AOL acquires mobile ad network Third Screen Media on 5/17/2007.
  • WPP Group acquired 24/7 RealMedia for $650M on 5/17/2007.

A few billion here, a few billion there, pretty soon you're talking real money.

What's happening is different for each player, but the overall trend is the same - expanding beyond paid listings into creative branding. Paid search was $6.7B last year and brand advertising was $3.3B. This interest in brand advertising may be a reaction to the expectation that television - which is mostly branding style ads - is moving online.

In Google's case, they are buying a company that has been successful with creative ads, essentially banner ads. Banner ads are the most common choice for branding rather than being used for actionable ad listings. They also inherit distribution agreements for AOL and MySpace, and more distribution capacity helps draw advertisers into Google to bid for placement.

Microsoft hasn't done well in any online ad segment - listings or branding - and with their acquisition they will be more involved with holistic ad campaigns and deep in the creative arena of advertising. The includes the ad agencies that do the actual construction of creative ads and create very innovative branding experiences like custom website which are blurring the lines between interaction and advertising. There may be some future tie-in with 'rich internet applications' and Silverlight. I may be wrong, but I don't see aQuantive providing additional distribution capacity - 'inventory' as the ad industry calls it. They of course claim it 'extends their platform'. Everything's a platform to Microsoft. Maybe they should simply try providing value instead.

The hope is that the contextual and behavioral profiling that is done for ad listings will be applied to better target brand advertisements. A requirement for this to work is for the ad network to know a lot about their audience - something that a single site cannot accomplish. Effective audience profiling is orthogonal to Web sites - orthogonal to the Web's organization. However, by looking at the architecture and technologies of the Web you can see the areas where this multi-site capability can exist :
  • clients such as rich internet applications, browsers and browser extensions like toolbars
  • intermediaries such as the proxies that ISPs like Comcast operate
  • compound resource structure of current Web documents. Since each resource can be retrieved from different domains, information can leak between domains.
Look for control or partnerships in these areas in the future.

May 14, 2007

Command Links

Uh oh. I see a mess coming up...
From Jakob Nielsen's Alertbox post on
Command Links
"Windows Vista introduced a new GUI widget for commands: the command link. Once something is in the system that people use on a daily basis, it becomes a de facto standard. Because they'll encounter them frequently in Vista, users will come to know and expect command links."


From what I gather, Vista 'command links' appear to be glorified buttons for native applications, not 'underlined text' links on Web pages.

However, Jakob continues with the following regarding web page links:
To reduce confusion, link text should explicitly state that it leads to an action and not just to a new page. It's not enough to communicate this info in the surrounding text; users often scan Web pages for the areas they can act on. Thus, you should assume that most users will only read the link text. In fact, users often read only the link text's first few words, so it's important to start with a word (typically a verb) that indicates the action that results if they click the link.


It appears Jakob is also implicitly approving the use of a link on a web page as an 'action'. From my reading, I think this is seriously wrong. I believe web page 'action links' need to satisfy the following two requirements

  • visually distinct from normal 'safe' links
  • syntactically distinct from normal 'safe' links


The second point is important because of the large number of automated agents that traverse the Web through hyperlinks. Adding unsafe links into the web will cause confusion among both people and software agents.

May 09, 2007

It’s my Vineyard

I found this blog about a couple that have purchased a vineyard in France - what a dream!

It’s my Vineyard
"Sitting in glorious sunshine on the terrace of The Restaurant du Pont having a delicious lunch we are reflecting that it is almost two years ago that we fell in love and bought Maison des Bulliats and its vines in Regnie, Beaujolais, one of the most idylic spots on earth."


This reminds me of the book "The Olive Season" which is about an impetuous British actress that purchases a run down villa and Olive 'garden' in the South of France. We found that book in the apartment we rented while we stayed in France two summers ago. Interesting book - although the scattered personality and writing of the author can get exasperating - and I'm looking forward to following this French vineyard blog.

May 08, 2007

XML is in the House

I ran across this page - Legislative Documents in XML at the United States House of Representatives: "The purpose of this website is to provide information about the ongoing work of the U.S. House of Representatives in relation to the eXtensible Markup Language" - and thought the host name was very cool - xml.house.gov.
This page has pointers to DTDs and schemas for the governmental processes involved with making and amending laws. So the next time we need to form a more perfect union, we've got this going for us.

And that page led to this page of a summary of floor proceedings of the US House of Representatives. It reads like a blog, only with less detail and there is no RSS feed.

May 03, 2007

Apollo, Silverlight, blah blah blah

Looks like Hugh W is one of the few questioning the value of Silverlight and RIA for the Web.

Flash, applets, Silverlight, Javascript -- the more you use them, the suckier your web apps are at exploring the web information space. I don't think it has to be this way, but it takes a design discipline few seem to have. These programming models are from the 80s. They have web APIs, but they're not web oriented. Programs end up as little desktop applications, not web apps. I don't see Silverlight changing that. It is good to have super expressive widgets -- hear hear. But if you're not pushing a bunch of hypertext down to my browser, you're not helping me explore the space.


I agree with his sentiment - you might say that RIA is to user interfaces as RPC is to messaging interfaces : more is not better. There probably will be a few years of smooth looking but hard to use (and harder to re-use) applications, while we wait for people to re-learn the basics of usability. I can only hope folks read Nielsen's UseIt column.

May 02, 2007

Mathematica 6

It looks like Wolfram Research has release a huge new release of Mathematica. I have only dabbled in Mathematica but I could spend all day playing and learning with it. It's hard to believe Mathematica came out almost twenty years ago!

Something they've added recently - and apparently improved on - is server based computing and visualization. Imagine what could be done by putting this together with something like Amazon's Elastic Compute Cloud (EC2).

See the Wolfram Blog for more details.

May 01, 2007

REST - it's inevitable

A couple weeks ago I was having lunch with a friend at Amazon - who coincidentally used to work at MS with the data access team and is a brilliant architect and engineer - and we talked about REST and how he was helping use REST concepts in refactoring some core back-end services. I made the observation that I was never stressed about how long it has taken for REST to achieve common industry understanding and acceptance - I just said "It's inevitable". I didn't realize just how short it would be for inevitable to show up.

It looks like REST has taken Redmond by storm. First I read that Mark Baker did some consulting with Microsoft (I missed the chance to have dinner with him when he was in town due to email snafu - major bummer). Then I read Dare Obasanjo's post that says "REST is totally sweeping Microsoft."

We passed the tipping point quite a while back, but it's still good to see pragmatic architectural sensibilities take root finally.

April 25, 2007

Sparkly things factory

Oh, this quote from Performancing.com is beautiful!

"Snap's preview anywhere gizmo is ruining the reading experience for millions of people. Its intrusive, obstructive and unuseful in almost every respect and use case. The fact that so many big blogs are using it, big well respected blogs, does not mean that it's useful, it just means that they, like most bloggers, have all the self restraint of a magpie in a sparkly things factory."

April 24, 2007

Amazon.com Widgets and hypertext

It looks like Typepad and Amazon are collaborating to help bloggers link to Amazon products. They have three widgets outlined, and one of them - the quick linker widget - uses custom attributes on an HTML anchor tag to make it easier to reference a set of products. The 'old school' would just use a URI, but those are hard to construct, hard to type and the wizards slow people down, blah blah blah. This not-quite-micro-formats approach is more understandable and more forgiving for hand-crafted markup. They define a new 'type' attribute for the anchor with the value "amzn". Then there are several more attributes like 'search' or 'category' or you can create a direct link via an 'asin' attribute. I assume some snippet of javascript would scan the page after it was loaded and construct the URI on the fly and set the href attribute on these anchors.
This is a creative solution to the "how do you construct a URI" problem, but it does leave spiders out in the cold, breaking the hyperlinking which defines the Web.

However, there is a simple approach they could use which would make a fully declarative and locally described document work and continue to allow auto-discover of hyperlinks to work - add a 'meta' tag to the head of the document with the URI template that corresponds to the 'type' attribute. I think the URI template proposal may need to do a bit of work related to optional or conditional patterns, or the URI template could stay static and the type="amzn" could change to be type="amzn-direct-link" or some other more qualified value.

April 21, 2007

Chocolate. Real chocolate.

The other day I was lucky to tag along with Jordan to the ESIF (Early Stage Investment Forum) in Seattle. The event was held by NWEN (Northwest Entrepreneur Network) and is sort of a 'graduation' milestone for local entrepreneurs and startups. We weren't the only software startup, but there were many other types of companies represented - bio-tech, high tech, textiles and gourmet foods. Not just any gourmet food - gourmet, hand-crafted and wonderfully made chocolate. Real chocolate. Not the burn-your-throat, over sugared, waxy poo that comes from Pennsylvania. I had heard rumors of a coconut curry chocolate so during a lull in booth duty I wandered around to score some chocolate. Several years ago a friend visited Mexico and brought home some chocolate that was flavored with chili powder and granulated sugar and the taste was fantastic so I assumed this curry chocolate would be just as good. I found the booth with Theo Chocolate and talked with two very friendly folks and asked about the curry chocolate. Naturally they offered a taste and I must say it was great - spicy but smooth and rich with chocolate. I told them that I'm always on the lookout for new chocolate because my wife really likes good chocolate, especially Dutch dark chocolate. Suprisingly, they started pulling out bars and explaining each - this is 91% cacao, this is 65% cacao and so on. Their flavors are intriguing, not only do they have coconut curry chocolate, but also Chai Tea chocolate and exotic chocolate using cacao beans from Ghana and Madagascar and Venezuela. I have only tried the Venezuela chocolate (91% cacao) and it's really becoming addictive - crisp chocolate with a dusky bitterness without being unpleasantly bitter. I hope the 'Special Limited Edition' label doesn't mean it will be hard to get in the future.



(Oh, and I just found this useful Flickr tool to upload photos directly from your desktop with a right-click menu. Quite a time saver.)

April 20, 2007

Tim Bray stole my flowers

I regularly read Tim Bray's posts, but admit that I skip the details of a few posts about hardware now and again... however his recent flurry of spring flowers made me laugh because it's like he's talking about my backyard.
It first started with a star magnolia, his was blooming a bit after ours had started flowering - here's a photo of ours.



Then it was a rhododendron, which is fairly common in the Pacific Northwest. We have a white with pink tinge and half a dozen more that haven't bloomed yet - they come out at different times of the year each with different colors, very cheerful.


Then there was the trillium which is my favorite flower. My wife bought a bulb and planted it out back by a little pond because she knew I loved them so much. Three came up this year and I missed the chance to take a good photo, but Tim's will do very nicely.


And there are others - tulips, more rhodies, daffodils. The grape hyacinths are pretty, but their fragrance is the best. Put these near your front door and you'll really know when Spring arrives. You might need to check the particular kind of grape hyacinth - ours looks a bit different.

If Tim next publishes photos of hibiscus and sunflowers, I'm going to get suspicious.

April 17, 2007

CSS - new style, same old sheet

I've been working heavily in the world of HTML and browsers over the past six months and for the most part enjoy it. It's much much better than the scene six years ago when I was at KnowNow writing fairly advanced Javascript client code that was supposed to work on Netscape Navigator and Internet Explorer - what a nightmare. Today's browser world is infinitely better. Now people only complain about box models being a pixel or two off. Well, and there's that z-index bug Microsoft hasn't fixed and probably doesn't even know about, even though everyone else does...

Anyway, my most recent learning experience has been with information extraction from Web pages - essentially extracting meaningful keywords from HTML. I must say, there's a lot of room for learning to take place here. There are several research papers I've found that are really educational, especially those that talk about extraction in the absence of a large body of other documents (corpus) to measure relevance.

As I was going through some experiments I realized that doing a decent job of extracting text from HTML requires knowledge of what 'markup' is and what the particular elements of HTML are defined to mean. Extracting meaningful phrases from markup means to ignore the markup and get to the underlying text which was marked up. But then I began to notice something - in all the advanced HTML pages that use the latest CSS to accomplish 'semantic HTML' (a phrase that I've heard tossed around pretty loosely) something is going wrong. The underlying text that is marked up is becoming gibberish. This is due to the use of CSS for layout and ignoring the effect of the tags on the text. For example, when a span tag is applied to text it is considered an 'inline' element - the underlying text is not meant to be fragmented and split apart and any extraction tool (especially a naive one that I was experimenting with) should merge the text fragments before, within and after the span element with no whitespace. But often designers will add layout and margins to the span tag in order to visually separate the text - yet the underlying markup indicates the text fragments are contiguous. How annoying. There is a simple solution - tag the text as it is intended to be read and understand the difference between 'inline' and 'block' semantics for narrative text.

April 16, 2007

Inertial electrostatic fusion

I haven't heard any news recently about inertial electrostatic (confinement) fusion, so I was pleased to read a note (on the Google Research blog of all places) mentioning Robert Bussard (former Asst. Director of the AEC) talking about inertial electrostatic fusion. Apparently it's a video - I hate video. I'd rather read about it, but this should be good. I tried to get some information from EMC2 (Bussard's company) a while back, but they didn't answer emails. Hopefully this video will give some interesting info on what they've been doing.

Wow - this is very very cool. Bussard has been working for over ten years on Navy contract to research and build a spherically confined fusion reaction - and they've done it. The funding from the Navy dried up, his company has the patent and now the next step is to get funding to build a full scale prototype. After that - cheap and limitless energy for the world. Wow.

DoubleClick and Google

There's been a lot of industry angst over Google's pending acquisition of DoubleClick. I'm fairly new to the game of online advertising so I don't know enough to predict the fallout.
Here are two good posts to read - the first is from Pulse360 Blog and is a sky-is-falling viewpoint. The second is better and is from Jordan's blog (the CEO of the startup I work for) and gives a good analysis of the value to Google and why they made the acquisition.

There are a lot of terms that I wasn't familiar with several months ago, so here is my cheat sheet of terms:

  • banner advertising - wide advertisement usually at the top of a page
  • inventory - this is what web page publishers have to offer, the space on their pages and the audience that will view those pages
  • remnant inventory - areas of a website that are not very popular and the web page publisher cannot charge lots of money for ad placement (because there are no viewers)
  • ad network - provider of ad listings
  • eCPM - effective click-per-mille (which is click-per-thousand page views)
  • creative - the 'creative content' that shows up in the ad listing. Google made big bucks doing just text and links and left the annoying flashing ads to others.
  • AdSense network - the Google system that provides ad listings to folks publishing web pages. It's supposed to provide an ad listing that is highly relevant to the page it is injected into, but that sometimes doesn't work.


(more words here)

It's interesting that the space on web pages is called 'inventory' - coming from Amazon I had a different view of 'inventory'. There are a lot of recurring themes between the world of product catalogs and inventory management that I am familiar with and the new world of online advertising. Maybe it'll all make sense eventually.

April 13, 2007

Sinterklaas Boot

And now for something completely different...
Over on my music blog I posted about a great techno dance song called Boten Anna - but this Sinterklaas Boot video which is a spoof from Holland (those crafty Dutch!) is the real gem!

April 11, 2007

Space Conference

Here's something interesting from Wired Private Launches, New Tech ... This Isn't Your Parents' Space Age
This week an estimated 7,000 government officials, corporate representatives and space enthusiasts will converge at the annual National Space Symposium here to hash out the technological, cultural and political issues surrounding the next decade's push for manned exploration of space.


I hadn't heard about the Constellation program from NASA, but it's very exciting to see the money put into space resulting in good work.

Here are a couple space blogs with a little news Space Politics and Space Report.

April 08, 2007

How to (Teach how to) Write a Spelling Corrector

Through Bill de hÓra's Bzzt Questions blog post, I found this post from Peter Norvig on How to Write a Spelling Corrector. The spelling corrector post was interesting initially because I've been doing a little text processing recently and have his code echoed the simplicity of approach I needed to use to squeeze the algorithm into JavaScript for use in a widget. But it wasn't strictly the code that really struck me - it was the multi-faceted learning opportunity that the post represents. On one hand, there is the lesson that thinking well before coding is a Really Good Thing. In this case, thinking of the problem in terms of lists, sets and maps clarifies the tasks that the software should perform. On the other hand, demonstrating the simplicity and expressiveness of Python shows that the actual tool used can be important. Especially a tool that removes obstacles between theory and practice. And on the third hand, Peter Norvig's post is a great example of education. It educates readers about processing natural language in the wild, it educates programmers on how programming languages reflect the mental model of the developer, it educates designers on how theory can and should influence the practice of software development and at a higher level it educates everyone on what a real engineering looks like.

Oh, and check this out
Fortunately, Google has released a database of word counts for sequences of up to five word sequences, gathered from a corpus of a trillion words.

April 05, 2007

Spring in the Northwest

Spring has definitely arrived here in Kirkland. One of the rhododendron and the Star Magnolia tree are in bloom, brilliant white blossoms finally brightening the yard as well as adding their perfume to the air. There is something about the warming of the forest at the beginning of the year that gives a smell of the outdoors that charges me up. I'm very lucky to be able to work from my home over the past six months and I took a break from work this afternoon to split some wood from trees that had fallen in our winter storm. (If you ever get tired of spinning your wheels and want to do something with visible results, come on over, pick up an axe and split some wood!)
Being outside is so wonderful - the sound of birds (especially the Northern Flickers), the random bumbling path of a bumble bee and even the first butterfly I've seen this year all make me grateful for the place I live. There was even a pileated woodpecker hopping around the lower part of the Doug Fir outside my office window. If my cat hadn't been asleep on my lap he would have gone nuts!

March 24, 2007

SimpleBits

Although I spend most of my time working in and around server-side systems, once in a while I find myself designing and building client side software. Some of that work is the visual and interaction bits, which is more challenging than designing servers - at least for me. Over the past few months I've worked heavily with CSS as applied to HTML and XML. Using CSS is fantastic. Sometimes infuriatingly inconsistent, but fantastic. I have a fascination with good 'design' - the visual kind of design that I appreciate but can't quite create. One site I stumbled on recently is SimpleBits by Dan Cederholm. This seems like a good starting point for exploring the world of page design, styling, typography and user interaction.

March 19, 2007

SpaceX launching demo flight #2 today

I don't know where I've been, but I had no idea that SpaceX corporation is so far along with their technology and their business. They three configurations - light, medium and heavy. These can reach LEO, GEO and the heavy can support planetary missions! A commercial company that can reach another planet! They are broadcasting the launch also.

This video of their Falcon1 launch in late 2006 kicks so much ass over the BlueOrigin test launch - although when BlueOrigin launches that cute iPod of a rocket for real, that will be so very cool.

Update - the launch was delayed 45 minutes, but the live webcast is on. Cool! I can see vapors from the fuel topping off! "T minus 20 minutes and counting..." Wicked!

Update - AARggh! They aborted launch at T minus 1:15. Not sure why yet.

Update - The launch engineer has announced termination of the webcast, but the launch will be rescheduled for tomorrow same time.

(Check out the google maps of the Omelek Island launch site - someday we'll have live Google Earth with the 3D trajectory!)

PS - if you want to chat about this, download this Firefox sidebar and go to the SpaceX home page

February 27, 2007

The Younger Generation

So I was sitting at my computer when my daughter walks up and asks - "What are you working on?" I told her I was reading about Web technology and happened to see a page on Wikipedia about HTTP. Since I knew she had seen Wikipedia before, I thought I'd impress her by saying that I had added a sentence to this page. So she says - "Oh yeah! I added a sentence to this page here." and she proceeded to take over the mouse and navigate to the page on Cubism where she showed me the sentence she added there - the second sentence at the top of the page. I was the one impressed! And she's not yet a teenager...

February 14, 2007

Job openings at OthersOnline

I've been busy and having a lot of fun cleaning up and building out different pieces of the OthersOnline system. So far I've fixed a few issues in the InternetExplorer toolbar, implemented the Firefox toolbar, implemented a widget for bloggers, built Web pages for profile search and browse, and variously tweaked and tuned the profile matching algorithms on the server. But now it's time to add a few more engineers. We have two job openings for an experienced software engineer and an experienced UI designer. Feel free to contact me at mike at othersonline dot com, or forward this to someone that may be interested.

Senior User Interface Designer
As a Web applications developer with OthersOnline you will define and create the user experience for all our applications that results in a highly social and interactive community of online users.
You will be responsible for operational support as well as prototyping features and brainstorming ideas. You will design, document and implement usable, enjoyable and engaging features using DHTML, Javascript, CSS, XML and Ajax and work within a Linux and Java based server environment. Candidates must be creative, flexible, self-directed, and understand how to design and write elegant, responsive and maintainable user interfaces.

Senior Software Engineer
As a software engineer with OthersOnline, you will create innovative algorithms and data services to dramatically improve search result relevance and create highly scalable, reliable Web services that extend the reach of our platform.

You will be responsible for operational support as well as brainstorming features and ideas. You will design, document and implement scalable, flexible features using Java, XML, SQL and HTTP in a Linux and Oracle environment. The ability to document technical approaches completely, correctly and concisely is required. Candidates must be creative, flexible, self-directed, and understand how to design and write high-performance, reliable and maintainable code and data models.

January 21, 2007

Social Content Services and REST

This article from Leigh Dodds about Connecting Social Content Services is an extremely well written summary of several social content services, a greatly pragmatic description of REST and an extremely useful review of the service APIs from those social content services.

If you need a quick list of aspects for engineers to consider when building a ReST based system (the list is derived from Joe Gregorio's site), try this:
  • How does the API use URLs to identify resources?
  • What HTTP methods does the API support?
  • What status codes are returned?
  • How are users and applications authenticated to the API?
  • Are resources linked to one another (use of hypermedia)?
  • What data formats does the API support?
This review notes that nearly all services do not use hypermedia which I think is unfortunate but understandable. I've always had a problem resolving the desire to be flexible in allowing the internal data identifiers to be used in many situations and the desire to be trivally easy for clients to access other resources by simply using links - the mashup problem. One issue I have is that the server-side software that generates the representation might not know all the possible resources made available by sibling services. Think of a US postal zip-code - if you have a service that provides weather based on zip code, should that representation also be responsible for linking to all other services - either provided by your system or some other server - that could potentially take in a zip-code? My approach is to return both direct links to known resources (tagged appropriately) as well as the short-form of the identifier, the plain zip-code for example. Microformats sort of do this, but it's a style that isn't well applied by data services.

One comment I have on this review of RESTful services is the author mistakes "URL" for only the portion before the query terms. This results in judging these services to be less RESTful than they actually are. Use of query terms does not mean there is one 'controller' with parameters, it actually increases the number of addressable resources which is good.

January 17, 2007

Funny story about a bank

Want to hear a funny story?
Okay.
So, this guy walks into a bank south of Seattle, WA.
"Hi" he says "my name is Mike Dierken".
"Here's my social security card, bank account number and Utah driver's license." He hands over a check for $5,000 and wants to withdraw $4,000.
The bank says, "Duh, okay. Sounds good to me." and hands over $4,000 in cash.

Here's the funny part…
The next day, same guy walks into a different bank.
"Hi" he says "my name is Mike Dierken".
"Here's my social security card, bank account number and Utah driver's license." He hands over a check for $5,000 and wants to withdraw $4,000.
This time the bank teller calls the other Mike Dierken who says "What the #!**?" while hearing the teller saying "Sir, please just have a seat over there. No, don't leave…"

I'll be out of the office tomorrow morning. I have a meeting with my local Bank of America branch. I must say, the teller that called me was very smooth and professional, pretending to talk to a 'signature authority' while talking to me. She made her side of the conversation sound normal while still responding to my questions. I imagine even check fraud could be frightening to a teller, as you don't know how they will respond.

January 08, 2007

Second Life client under GPL

I've never used Second Life, but I like the idea of people creating hyper-spaces. I used to do work with 3D graphics and have always thought the multi-player online games had poor visuals - the shading (coloring of models) and lighting really need work to be realistic. I'd like to see 'scenes' with near movie quality lighting and shading that you can invite a few friends into, yet hyper-link to other spaces. I've even toyed with the idea of a 'better' markup language for 3D - less modeling and more auto-layout capability.

Anyway, I don't know what it will mean to have their client library freely available, but this post on BoingBoing had a comment that really struck me:
"Customers only ever get to love it or leave it. Citizens get to change it."

January 05, 2007

Happy Birthday Grandma!

Today is my Grandmother's 90th birthday. Happy Birthday!
She lives in Montana, by the Little Blackfoot River just a few houses down from my Uncle. It's always great to go visit. Unfortunately we didn't fly out to see her for her birthday, airline tickets were just too expensive. We'll probably drive out in the Spring - at least the snow should be gone by then!

December 21, 2006

Windstorm 2006


Taking the limbs off.
Originally uploaded by dierken.
The recent windstorm here in Washington knocked down three trees in our backyard. One snapped at the base and landed on our swingset and narrowly missed our deck and hot tub. The other two tipped over, pulling up a huge rootball and landed on the edge of the yard and into the forest.
We were without power for five days, which wasn't so bad for us. Our fireplace in the basement kept us warm and we learned you can actually bake cookies on a barbecue grill (if your daughter really really wants to make cookies). Some friends dropped their kids off for a sleepover (since our house wasn't freezing) and we all had a good time playing games by candlelight.

Now it's time to finish taking the limbs off the trees and clean the yard - fun fun.

December 14, 2006

Out of Amazon

Today is my last day as an official Amazon employee. I've been on leave of absence for three months and have decided not to return at this time.

This was not an easy decision - recently I was talking with a team at Amazon about a project they have started and it's exactly the kind of thing I've been wanting to build for years - but there are some other things I'd like to work on elsewhere.

I think Amazon is great, they have smart people and are in a great position to grow the Web. I expect to continue to see great things launched by the people at Amazon and really, really look forward to the project my friends are building.

Thanks for everything!

December 05, 2006

Happy Sinterklaas and St. Nicholas Day

Happy Sinterklaas and St. Nicholas day!
When I was little, every December 5th was the day we would wake up and find a little package or an orange in our shoes by the fireplace. This is a family tradition from my mother, who grew up in Germany and taught us about St. Nicholas.
And since my wife's family is Dutch, we also get to celebrate Sinterklaas on the 5th - today we opened a package from Sinterklaas and Zwarte Pieten (Black Pete) with my favorite ginger spice cookies and we each found a giant chocolate letter - the first letter of our name. I think there weren't any 'M' letters left, so I got an upside down W!

December 03, 2006

Why I like blogs

This post over at
edgeio is why I like blogs. It's real, it's about people and and it's about what people are doing. In this case, we hear a brief overview of their first rollout of real-time searching of a hundred million offer listings.
Cool stuff.

November 28, 2006

Agents, toolbars and adware

Here is a thought provoking post from Brian Smith about agents, adware and toolbars on ComparisonEngines.com. I like that he presents the core value of 'adware' and phrases it in a more general context of 'agents'.

Long ago I had been very excited about agents and event driven applications but spent most of the recent past working just on messaging systems (subscriptions, notifications, etc) rather than the agents themselves. Maybe it's time to go back and revisit agents. Recently I've had this notion that an interesting paradigm for empowering people to build agents might be to use a spreadsheet with event inputs and message outputs - but I haven't even started down the road of working out what it would look like. I think general folks wouldn't understand specialized agent building languages but might be able to poke things together with a visual tool.

November 20, 2006

The Spreadsheet as Mashup Fabric

The title of Joe Gregorio's recent post - The Spreadsheet as Mashup Fabric - really caught my eye. For the past two days I've been thinking about the spreadsheet as a simplifying tool for building event driven web applications - accepting events, summarizing many messages, routing and republishing the results to a new set of subscribers.

I've just started thinking about this so I don't have any actual situations it might be useful for, but the idea of opening up a compact, rich and information dense document in a very responsive interactive application for ad-hoc applications and messaging components sounds interesting.

Sort of like message-driven beans on acid.

November 17, 2006

What's old is new again

Check out this short article from 1998 on making more responsive web applications using - get this - javascript.

We can create a more responsive experience if we look at the available technology in a different way. In particular, we are going to forget the assumption that a Web page is loaded directly from a Web server. Instead, we are going to use the request/response structure of the Web in a slightly different way.


I wonder how far back we can see this approach being used or clearly talked about.

November 14, 2006

37Signals and Google Web Accelerator

From the rest-discuss list, here's a post from 37 Signals that missing the point about the Google Web Accelerator fetching URLs.

This wouldn’t be much of a problem on the public web since it’s pretty tough to be destructive on public web pages, but web apps, with their admin links here and there, can be considerably damaged. If you have a web app, it might be worth returning a 403 when the HTTP_X_MOZ is set to “prefetch” header is sent. This will keep Web Accelerator from clicking destructive links.


I like this "destructive links". A major point of the Web is that links are not destructive.

I think the better approach is to not use GET to modify or remove data. That's simply unnecessary and against the word and spirit of HTTP.

Remember, a link is not a widget. You can't use a simple anchor element as if it were the same as a push button. Whenever an HTML page has an anchor tag, it is explicitly advertising that the URL is safe to retrieve. By sprinkling your HTML with these landmines, it's the author that made the mistake, not the browser or web accelerator.

If you are writing web apps, it might be worth more to think about what's happening and why, rather than hack a workaround that only occasionally avoids a bug introduced by your own application. Fix the bug, and no workaround is necessary. Again - a link is not a widget.

November 08, 2006

New Direction is new theme for Democratic plan

Is it just me or does the phrase
'New Direction' sound just like 'nude erection'?

I have a bad feeling about this...

November 06, 2006

Participation Inequality: Lurkers vs. Contributors in Internet Communities

A few years ago I had stopped reading Jakob Nielsen's Alertbox even though I had been reading it for a long time before then and had enjoyed his fact filled articles and analysis. Just the other day I popped back in to check out his recent postings and this one caught my eye - Participation Inequality: Lurkers vs. Contributors in Internet Communities - as I'm looking into communities, social aspects of the Web and such.

The article gives some history and data that covers the basics of a 90-9-1 rule with respect to participation - most people just lurk, some interact occasionally and a few are very active. Many people know this and if you've read Clay Shirky's "Power Laws, Weblogs and Inequality" article you'll be familiar with the idea. Although there have been comments elsewhere that Jakob doesn't "get" RSS (and his alertbox site doesn't have an RSS feed), I think Jakob does understand the implications of RSS and this post provides some value by highlighting some issues and suggesting how to design useful systems in the context of participation inequality.

Search. Search engine results pages (SERP) are mainly sorted based on how many other sites link to each destination. When 0.1% of users do most of the linking, we risk having search relevance get ever more out of whack with what's useful for the remaining 99.9% of users. Search engines need to rely more on behavioral data gathered across samples that better represent users, which is why they are building Internet access services.


Make participation a side effect. Even better, let users participate with zero effort by making their contributions a side effect of something else they're doing. For example, Amazon's "people who bought this book, bought these other books" recommendations are a side effect of people buying books. You don't have to do anything special to have your book preferences entered into the system. Will Hill coined the term read wear for this type of effect: the simple activity of reading (or using) something will "wear" it down and thus leave its marks -- just like a cookbook will automatically fall open to the recipe you prepare the most.



(Bummer - Jakob Nielsen had a 'User Experience 2006' in Seattle a few weeks ago and I missed it...but maybe I can catch some of it on a blog somewhere)

Added: Here is another good article from Jakob Nielsen about traffic log analysis.
The top 10 queries accounted for 10% of the total traffic, so each one of these queries is obviously more important than those that brought only one visitor. Taken together, however, the single-use queries accounted for three times as much traffic as the top 10 queries. This statistic shows the folly of focusing search engine optimization solely on a few high-performing queries. Your site must be found when users enter relevant queries -- and the possibilities are typically vast.

November 05, 2006

Abundance .vs. Scarcity

This is a good post by Dare Obasanjo commenting on economies of abundance .vs. scarcity with respect to the digital abundance of Internet and the scarcity of time and attention with respect to human intelligence. The Economy of Abundance and Other Fairy Tales.
I first heard the phrase 'economoy of abundance' many years ago in a story called "Manna" by Lee Correy (pseudonym of G. Harry Stine) - a neat science fiction story. I wonder where it first came into use...

Most successful Web companies today are exploiting the scarcity of attention and time that plagues all humans. In a world where there a hundred million websites the problem isn't lack of content, it is finding the right content.

October 11, 2006

WS-Notification approved by OASIS

And in unrelated news, OASIS members have approved WS-Notification v1.3

OASIS, the international standards consortium, today announced that its members have approved WS-Notification version 1.3 as an OASIS Standard, a status that signifies the highest level of ratification. WS-Notification defines a pattern-based approach for disseminating information amongst Web services.


I haven't read the relevant specifications, but should make the effort. Nah... I'll just go for a bike ride in the forest instead.

Benjamin's REST Tutorial

This is a really good engineering document that gives pragmatic design and implementation guidance for building RESTful systems: Benjamins REST Tutorial:
My favorite quote:
"You are starting to see network effects as programs with similar schemas and content types can 'just talk to each other'. They are running out of things to disagree on, and are running out of software that needs to be written to support the differences."

September 13, 2006

Taking a break from Amazon

I've decided to take a break from Amazon - tomorrow is my last day. Officially I'm taking time off as a leave of absence, but at this point I'm not sure of the odds of returning full time. However, I have talked with a few groups and there is some mighty interesting work coming up - one possibility in the Web Services group and another in a 'new business' sort of group.

The past three and a half years have been great and full of action, but that action has me pretty much burned out. I was really, really lucky to start in the group that I did - the product catalog group. After a couple years I wound up managing the team that is responsible for the primary storage of everything offered for sale on Amazon's platform - a couple hundred million listings. The reason I enjoyed this area is that while there are many things that Amazon does, one of the core areas is allowing sellers to offer thing for sale. There is both the satisfaction of scaling to huge numbers, and the satisfaction of enabling the 'little guy' to list stuff and make some money, possibly even make a small business. There are so many people that created a business or run their business based mostly on Amazon that it's both humbling and humanizing - believe me, we all take it very seriously.

The people I've interacted with at Amazon are the best I've ever worked with, and I've been at several companies over the past fifteen years. The really amazing thing about Amazon is that everyone is uniformly smart - not just engineering people, but everyone. That alone makes things so much easier.

August 16, 2006

All the Data. All the Time

Remember when nobody had a web server, commodity hardware, gigabit ethernet or scalable storage? There was (and still is) a huge industry building special purpose computing machines for special application - CAD, electronic design, video and so on.

Then some innovative folks realized that access to a hard drive would become slower than access over a network and the slide away from direct attached storage to network attached storage began.

Directly or indirectly, that shift was echoed in this year's launch of Amazon's S3 web service - storage on a dime (well, on fifteen cents).

Today I had an amazing suprise when I got home from work. My brother has been preparing to paint the exterior of our house and while in the area he goes bargain hunting, both on Craigslist and Live Expo (I'll refrain from my normal rant about local listings at this point.) In the back of his pickup was a huge black cabinet - an Auspex NetServer 2000. This is a 700GB storage server that was wicked cool back in the day. It could have two additional storage nodes attached to get up to 4.5TB - all for the measly sum of $70,000. And it's in my brothers pickup. And he snagged it for free. For free. Of course, it took three people to heft it into the truck and we have absolutely no idea what to do with it, but my gawd this is cool.

Among available NAS designs, only the Auspex 4Front NS2000 (Auspex NetServer 2000) series of content servers uses both approaches, making it the most advanced NAS product design available. In this design, each of three I/O nodes has two Intel processors. Each processor runs specialized real-time software called the DataXpress kernel. One I/O node processor, called the Network Processor (NP), manages highly reliable customized software that controls all network protocol and caching functions. The other I/O node processor, called the File and Storage Processor or FSP, handles file system processing and storage processing. This design, called Functional Multiprocessing (FMP), allows network processing for a second I/O to occur in parallel with data retrieval for the first I/O. The FMP design provides advantages over both nonparallel single processor and Symmetric Multipro-cessing (SMP) designs that handle I/O activity in serial.


This whole discussion of a processing node sounds very cool - kernels, multi-processors, parallel operations - very advanced.

If anybody wants some of that high tech magic, let me know because I have a couple hundred pounds of it in a pickup in my driveway.

July 28, 2006

Norm on Names

From a post from Mark Baker, there's a fine post from Norm Walsh on names that you all should read - my favorite quote:

Contrary to what you may believe, there is nothing about the “http” URI scheme that requires use of the “http” protocol. And even where the “http” protocol is used, there's nothing about it that requires access to any particular machine.

On your desktop, your web browser may return things from its cache without ever hitting the web.

July 02, 2006

Erlang - the next buzz in web conversations

This post by Patrick Logan points to another post talking about Erlang and scalability. I've been hearing a lot about Erlang from the REST community recently and I've been thinking about all the things I could build with yet another web server that supposedly supports over 60,000 concurrent connections.

July 01, 2006

Rock 'Em Sock 'Em Robots


Ah, the good old days of Rock 'Em Sock 'Em Robots. When I was a kid this was a favorite game and it looks like Mattel has a new re-issue of the game.

The interesting thing about this particular page on Amazon is that this game is offered by three separate sellers, and Amazon is one of them. Before today, this was not possible due to the exclusivity contract Amazon was operating under. My team was only peripherally involved with the effort to make this happen, but watching the speed and professionalism of all the teams drilling into all the details, crossing the T's and dotting the I's, really made me proud. Not only did the people here work quickly and competently, but the software platform worked as designed.

Another interesting thing about that page, the customer reviews have multiple categories of stars - durability, fun, educational and overall - I never noticed that. Those are exactly the areas I look into when reviewing toys and games. Pretty cool.

Let the Rock 'Em Sock 'Em Robots competition begin.

June 29, 2006

RESTful Interface Description

The other day, a friend from Oz asked about IDL (interface definition language) for RESTful services (it turns out WSDL 2.0 has binding definitions that support all HTTP methods) and now there is this post from Phil Windley Crying Out for a RESTful Service Interface Description Language:
The only way that we’ll get to a place where Web 2.0 apps are more easily integrated is when we have a service interface description language and other metadata standards for RESTful services.


At the top of the post he links to another blog by Dave Rosenberg that says in part "To me the big opportunity of Web 2.0 development is the ability to create a better user experience based on features etc."

Following Phil's post is a comment starting with "What we need is a [...]".

Here's a suggestion - pick two services you would like to integrate, something that would result in real tangible value that you would actually use day-to-day, and try actually implementing these difficult integrations (and some are tricky) and write about the effort and the problems. That real effort with real implementation will surface the real problems that need solving. Crying out "you should do this" or "people need that" is just so much wishful theory talking.

In theory there is no difference between theory and practice. In practice there is.

June 28, 2006

The Web Is a Pipe

From The Web Is a Pipe:

Opportunity opens up when we use HTTP to connect our server infrastructure components together. If you use HTTP between your front-end web server and your back-end application server, suddenly you gain the ability to swap Apache for Lighttpd on demand without worrying about the FastCGI bits. Your application server won’t notice the difference. You can replace either Apache or Lighttpd with Pound or Pen. You can even replace them with some sort of hardware load balancer solution if that floats your boat.

Even better, when you use HTTP as the glue, you suddenly can use a whole host of tools to probe the various parts of your application. You can use curl to probe just your application server. Or even point a browser at it, assuming that you have a clear path through your firewalls and what not. And, if you’re really l33t, you can do a manual telnet and make just the request you need—with all the right headers—to simulate exactly a problematic client request.


Exactly. Using a web server as your application server means you can replace your application server with yet another web server.

June 25, 2006

HDR Coconut


big coconut
Originally uploaded by Haiku Garry.

I stumbled across this surreal photo of a coconut and was intrigued by the tag "HDR". What is "HDR" I thought? It turns out that HDR means high dynamic range. This is more than merely 32bits per color channel per pixel, it usually means capturing the full dynamic range. Most cameras can't do that, so the trick is to take three or more exposures and blend them together - called tone mapping or exposure mapping. Flickr is full of amazingly artistic photos with rich and deep colors. Many looks not quite right, almost as if they are a bit - but only a little bit - beyond current capabilities of computer graphics lighting models.
The HDR pools on Flickr have a lot of abandoned heavy iron cars from America's motor past, many photos of solitary houses, of disorderly mechanical/industrial buildings and one photo stream of an apparently abandoned castle.

June 20, 2006

Kaboodle

I found Kaboodle through a Google search alert (for 'social network shopping'), and for some reason I think the idea is really cool. It is sort of similar to Ta-Da lists from 37signals but more of a rich media list maker - wishlists, compare product data from different sites (very useful when hunting for a digital camera), assembling details on places to go for vacation (I could have used this last year...). Lots of uses come to mind.

The color scheme is orange and blue, like all good startups, but the site still looks fresh.

June 15, 2006

S3 support Virtual Hosting

This is good news - S3 supports the Host header in HTTP requests. I had been meaning to write about the security holes is storing different people's data within the same domain - as soon as two people host javascript, then 'cross site scripting' becomes possible within one host domain. This enhancement allows folks to trivially avoid that problem (assuming people want to host HTML and Javascript on S3 - not it's advertised purpose).

June 10, 2006

ACM Interview with Werner Vogels

I finally found time to read the ACM interview with Werner Vogels about Amazon and service oriented architecture (conducted by Jim Grey, no less). He talks a bit about Amazon history and a bit about how services affect development teams as well as the runtime and operational benefits.

Here's one of several 'lessons learned' that Werner recounts:
A second lesson is probably that by prohibiting direct database access by clients, you can make scaling and reliability improvements to your service state without involving your clients.


Interesting - since a database is also a service, what is the essential difference between direct 'database' access and direct 'service' access that improves the situation?

Another interesting comment:
Other lessons are related to how you access services: If you want to be able to aggregate services easily, if you want to insert advanced infrastructure techniques such as decentralized request routing or distributed request tracking, you need a single unified service-access mechanism.


Hmmm, the single unified access mechanism sounds like a protocol. I'm in the middle of building a new service and haven't been considering this - I wonder if I'm going to get in trouble. But I guess that's okay, since my business card does say Sr. Troublemaker...


Now this is the part that I like - as it involves my stuff!
About a million small and larger businesses sell on the Amazon platform. For example, if you go to one of the book pages, you will find that item is also available new or used from some of our many partners. These can be very small independent bookshops or larger retail operators that want to sell on our platform.


And onto REST...
Do we see that customers who develop applications using AWS care about REST or SOAP? Absolutely not! A small group of REST evangelists continue to use the Amazon Web Services numbers to drive that distinction, but we find that developers really just want to build their applications using the easiest toolkit they can find. They are not interested in what goes on the wire or how request URLs get constructed; they just want to build their applications.

May 21, 2006

Amber stones in the Northwest


Amber stones
Originally uploaded by dierken.

I've always been fascinated with minerals and gems, so this last Thursday I was very excited to go on a school field trip with my son to hunt for amber near Tiger Mountain.
There is a geologigist - Geology Bob - that conducts field trips for groups and schools to hunt for rocks, crystals, fossils and other cool things. From what I could hear (my ears are plugged up from a cold), back in the 70's he discovered a field of amber stones here in western Washington, which is astounding - there are only five places in North America that have amber deposits. Because he has been involved in local geology for quite a long time, he has received a permit to go onto state park land to conduct these geology tours for educational purposes. We learned about how amber is formed - ancient tree sap sinks into a swamp, gets covered in sand and mud, is pushed down into the Earth a quarter mile for ten million years, and then somehow surfaces again. It appears there are two earthquake faults in the area that are pushing together and squeezing material up to the surface. This area has coal, fossils and suprisingly some amber. Digging is fairly easy and when the shool kids find the shiny clear flakes and pebbles, they just holler with excitement - it was very fun. The class was full of third and fourth graders, and some of them were already planning on selling chunks of coal and real amber for a milllllion dollars on eBay. I just love that wild enthusiasm. They weren't even very disappointed to discover they would be lucky to get one or two cents - they figured they would need a hundred pounds, or maybe even five hundred pounds...

April 28, 2006

Microsoft is building a Google cluster

From Greg Linden's blog, a quote from Microsoft:

"the people who could build a viable [Web] services infrastructure of scale are companies that have both the will and the capacity to invest staggering amounts of money."


Hmm. I suppose that would exclude Google, BitTorrent and the Web itself. Unless they meant "the people that need to catch up will need staggering amounts of money, otherwise they lose".

dojo.storage: Client-Side Storage

From Ajaxian, here is a post about
offline access and client-side storage via Dojo. We are getting close to the tipping point for disconnected use and client-side state - and with the strength of the REST buzz (I can't believe how many people are actually applying REST!) this really bodes well for the next few years of Internet-scale application development. If nothing else, at least it's yet another way to route around the damage that is the Win32 API.

April 24, 2006

Pubsub .vs. polling - again

This is a great read from Bob Wyman on Dave Winer's comments about polling compared to notification :
From As I May Think...: Dave Winer: Show me that mathematical proof!

Dave's lack of understanding of the issues related to scaling can be seen in the history of the weblogs.com site that he struggled to build and maintain for so long. That site takes "pings" from blogs and then consolidates them into tremendous "change lists" which must be polled. Essentially, this site converts an efficient push-based update notification system (pinging) into an inefficient polling based system. Weblogs.com, as Dave built it, didn't even support common methods like eTags or RFC3229+feed to improve polling efficiency and scaling. The result was that it simply didn't scale and was frequently incapable of providing the service levels that people expected. Only now that Verisign has taken over the site and dedicated much better engineering staff and much more hardware, has the weblogs.com service begun to be somewhat useful again. However, since weblogs.com is still based on the terribly inefficient polling of change-lists that Dave supports, it is still a far cry from being what it might be.


Ouch. But I did like the nod to KnowNow and mod-pubsub. Not sure if mod-pubsub is still alive and kicking though.

April 19, 2006

Best Google Image

The latest home page graphic from Google is my absolute favorite -

Delta feeds - RFC3229+feed

About a year and a half ago, Bob Wyman was instrumental in defining an approach to greatly reduce the load and bandwidth used by applications that polled for changes to RSS/Atom feeds.
The other week, he noted that Microsoft will support the RFC3229+feed approach as well - which is good.

The only problem I have with this approach is that I think it is simpler to use hyperlinks, and I haven't seen a real comparison between the two. Both approaches have the client application maintain state of what data was last retrieved, but using hyperlinks has more chances for pre-existing caching servers to work without modification. I think the Atom protocol has defined something like this, but I couldn't follow the email threads.

To use hyperlinks, the data returned in a feed would have a link to the 'next' (more recently changed) posts. The client would then follow that link, which would either be empty (and optionally have a cache-control header to indicate how long to wait before checking again) or have more data - along with another 'next' link. The client just keeps following the links. The client would have two URIs - the original, well-known location that new readers start from, and the changing one which is the set of data most recently retrieved by that particular client. The server decides what the 'next' link is and what it contains - the data would be very cachable across all clients, merely by checking the URI.
The downside of this approach is the need to put the link within the content of the response - or add a response header for that location if the content isn't easily extended.

I put together a sample application that shows how this works - this is a simple html/javascript chat client, this is a link to the list of messages.

April 12, 2006

Broken as Designed

Sith Obasanjo in Broken as Designed can't decide which way is away from the dark side -
"However I do think some Web/REST advocates need to look around and realize what's happening on the Web instead of arguing from an 'ideal' or 'theoretical' perspective."


The Web advocates need to realize what's happening on the Web??

Okay, we all know some resources have broken content-type headers when you retrieve them. Others don't. Some clients never use content-type correctly, others do and many use heuristics. That's okay. Just do your homework - and part of that is to read the TAG finding and use it where appropriate - and be a good web citizen. That's simply practical advice based on practice. The "don't use content-type" is a theory that is unproven in practice.

Update - it looks like Dare found breakage in Cookies as well, but after further investigation it turned out to be a problem with the data returned by the server. But if the bloggers using RSS can't correctly control this Cookie response header - would we have no choice but to drop support for that header as well, as suggested for the Content-Type header? I mean, the theory of Cookies is all well and good, but if you look around to what the major players like MSN are actually doing, who are we to stick with something broken as designed merely because of theory?

April 11, 2006

Persistent Search and OpenSearch

It looks likes there's more interest (again) on saved searches and search alerts -
unto.net - Persistent Search and OpenSearch

Hopefully things have changed since I last reviewed this space in 2004.

Content-Type is dead

What a simply stellar idea here - Hixie's Natural Log: Content-Type is dead - browsers were broken, server configs are broken, so there can't possibly be any reason to use this header. The browser can't use it, so nothing else should either.
I think it may be time to retire the Content-Type header, putting to sleep the myth that it is in any way authoritative, and instead have well-defined content-sniffing rules for Web content.


You shouldn't throw away the Content-Type header even if server configs aren't easily controllable by the author. Go ahead and do the right thing - use the document context and tags as a hint on how to handle the content, use the content-type along with the content itself. There's nothing wrong with applications retrieving the resources referenced by an img tag to assume that the retrieved content is an image.
The only arguments people may have is when there is little context available when retrieving content (no hypertext source document) and the retrieved content could be interpreted in several ways.

April 09, 2006

Telomolecular Nanocircles

This description of telomere technology from Telomolecular looks interesting:

Synthetic DNA Nanocircles are a biomedical nanotechnology invented by Dr. Eric Kool and colleagues of Stanford University. These nanometer-sized circular DNAs have been shown to elongate chromosomal telomeres in vitro. They consist of DNA bases arranged in a sequence that templates the lengthening of telomeres by repeated addition of new TTAGGG sequences. Nanocircles have shown promise in telomere elongation in human tissues. By combining nanocircles with new proprietary gene therapy and delivery technologies, Telomolecular believes that nanocircles might work efficiently in living animals.

April 06, 2006

Google Base Storms Into Europe

From webpronews.com:

"The search advertising company will have a lot of catching up to do in Europe, where Amazon has relationships with businesses like UK-based Marks and Spencer. Google may have to pitch something beyond its online capabilities, though, according to the FT report:
One big UK retailer with no online presence said on Wednesday that Google's retail offer would be of interest if the internet company could also arrange for distribution. This potentially huge task has raised doubts about the long-term business models of other online retailers such as Amazon.com.

Doubts over Amazon's distribution don't have much grounding in reality. The company has built up its distribution network over the past decade and does have some knowledge in the area. Google's expertise at distribution probably does not go beyond making a change in a router's access control list and opening up a website. "


I like that last bit - Google's expertise at distribution probably does not go beyond making a change in a router's access control list and opening up a website.

The article goes on to suggest that Google team up with UPS for the distribution center. I wonder if that would work.

Google Base - Attributes

From a post by Mark Baker, I started looking at Google Base - they have a lot of interesting features - easy creation, bulk upload, extensible attributes and a
predefined set of attributes.
I've only looked at the tab-delimited bulk upload, but it looks pretty reasonable so far.

For listing products, the process seems pretty easy - http://base.google.com/base/a/1214338/9143072239359280355 . Their payment system is also integrated, so Google now has a real offer listing system with decent search. I imagine the next step is to syndicate the content as xml, wrap the data in a skin and provide good indexed search.

April 01, 2006

Service Oriented What?

Maybe I'm misreading this, but I almost fell over laughing at the post on Service Oriented Enterprise

3. REST, as the principle foundation of the Web, has proven that it too can scale - but the areas where it has proven to scale the best are related to the movement of unstructured HTML documents using a constrained set of verbs.


The but part cracked me up - from my viewpoint, it's the movement of unstructed documents and using a constrained set of verbs that allow for the scaling demonstrated by the Web.

March 30, 2006

Naked Answers

I like what Werner wrote about the visit from Shel Israel and Robert Scoble. Heh, wish I'd been there. I must have missed the email...

Anyway, this particular part struck me as being very true :
Amazon is a long time pioneer in the space of involving their customers with our product. And we really listen to our customers; any Amazon employee who encounters an issue on a forum or weblog or at any other place is empowered to escalate those issues internally immediately until they get fixed. Customer feedback is essential for Amazon and we will use all effective means to get it.


A few times over the past few years I've brought issues to internal teams and each time they have been totally positive about working on them - often it's a 'user education' thing (usually I'm the one being educated), but when it's a larger issue I've seen teams honestly re-examine what they are doing and incorporate the feedback and real change really happens.

March 26, 2006

Ripple - Internet IOUs

From a post by Mark Baker, the
Ripple payment system looks facinating! This is very intriguing because two nights ago I was wondering about payment systems and whether banks support 'request for transfer' requests in addition to just transfer-in and transfer-out requests. I imagine the problem of spam with request-for-transfer-out would be a problem. Anyway, the Ripple system is a distributed network of trust for the exchange of digital IOUs. These can be exchanged for products, services or even cash - anything of value. It'll be interesting to see what happens if someone integrates this with Edgeio or Craigslist (think 'mashup'). I haven't finished reading the whole site, so there are lots of open questions around fraud and actual overhead of participation, but this is all very cool.

Their About page is a good place to start.
From their FAQ:

How do I cancel debt on the system once it's been settled?

When two connected account partners settle a Ripple obligation outside the system, for example, by using cash, the person receiving the settlement must record it in the system by returning the settled IOUs to their issuer. They do this by making a regular Ripple payment to the person who settled their debt with you. Your mutual account balance will be changed to reflect the fact that the debt has been settled. In effect, they are using the IOUs they hold on the Ripple system to purchase cash.


The Internet is evolving it's own system of governing - of the people, by the people, for the people. First it started with communication, then community and now commerce. It seems inevitable that it's members will become an independent government unto itself. It won't start out perfect. Behavior, laws and regulations will continue to adapt, just as historical governing regulations have changed over time. But this tells me that there is a reason why this technology stuff is important. It is the technology of these Internet based systems that help private citizens remain private and remain citizens - and keep on being free. All changes start with a small ripple.

March 18, 2006

Irish pub on St Patrick's Day

On most Friday's, our team goes out for a lunch in Seattle's International district - mostly Asian, Thai and Indian food. But this Friday as we were heading out, we remembered it was St Patrick's Day and started looking for an Irish pub. The ones we knew about were kind of far, so we decided on going to F.X. McRory's - I always thought of it as a whiskey bar, but figured it had some Irish-ness to it.
When we got there, the staff all had some sort of green hat or shamrock painted on their cheek, so we knew we found the right place. There is a huge dining hall off to one side with high ceilings, old wooden floors and super tall windows letting in the sun - it was very nice. But the best part was that our table was right next to a band playing drums and bagpipes! I love the sound of bagpipes, but dang they were loud . We noticed later that the musicians had ear plugs - cheaters.
The menu had stew, shepherd's pie and corned beef and cabbage sandwiches - so I got the sandwich. I'd never had corned beef before and I have to say it's good enough try again someday - especially given a tall glass of Guiness to go along with it. But I'll have to warn you, the little dish of what looks like cabbage isn't cabbage, so if you take a big bite (like I did) you'd better really like horseradish.

March 15, 2006

ArtRage 2

If you or someone or know has even the slightest interest in painting, either on canvas or the computer, go right now and see ArtRage 2 - it's a fantastic program that feels like real painting. I really need to get a pressure sensitive tablet...

Pandora - find more music

Wow, the Pandora site that
plays similar music based on innate musicality is way cool! It's based on something called the Music Genome project that categorized songs on 40 or more dimensions. Fancy fancy!

"Ever since we started the Music Genome Project, our friends would ask:

Can you help me discover more music that I'll like?

Those questions often evolved into great conversations. Each friend told us their favorite artists and songs, explored the music we suggested, gave us feedback, and we in turn made new suggestions. Everybody started joking that we were now their personal DJs.

We created Pandora so that we can have that same kind of conversation with you."

March 14, 2006

Amazon S3 - finally, a real REST service

The new S3 Web service has what I would consider a very well designed REST interface that is also well documented (all the Amazon Web services have decent documentation). The current set of Amazon Web services is starting to grow nicely.

The S3 service is advertised as a Web-scale storage system with a super simple API and apparently very low storage and bandwidth costs. Each resource stored can be up to 5GB and also has support for ACL based permissions - to access the ACL, just append "?acl" to the resource. That's pretty simple and effective.

March 10, 2006

Microformats and product data

I've been hanging out on the microformats.org discussion list related to product data - too bad the email list isn't linkable. Anyway, these are my latest thoughts:

I am also wondering where and when a microformat for products would be useful. I'm generally wondering when /microformats/ are useful. (see http://microformats.org/about/) It seems to me that general narrative text and general conversations in text sometimes have 'hard' data like people, places, dates, times, phone numbers, etc Marking these as /being/ people, places, dates, etc. makes it easier for software to process - but the underlying data below the markup is still readable and sensible to people. Product data isn't like this at all. It's an unregulated mass of named properties with little rhyme or reason - there is no need to gently superimpose tags without breaking the elegance of the underlying data because there simply isn't any underlying elegance.

What it sounds like may be useful is a simple way for people to talk about things and reference those things, but without needed to fill in the blanks for unnecessary detail. For example, if I talk about the new mountain bike that I bought and immediatly scratched up by tipping over, it would be cool if I could tag the phrase "new mountain bike" as a reference to a product - but I don't want to actually describe the composition of the bike, the kind of tires and suspension and whatnot. People don't do that, manufacturers do.

If we want to solve this for manufacturers, then just use XML. Then the next problem is how can software determine which attributes from which manufacturers mean the same thing? That's where a shared 'open product attribute' database would be useful.

Gosling down on PHP and Ruby

This is a hoot - from
James Gosling at a Sun conference - narrowly correct but broadly missing the point :

'There have been a number of language coming up lately,' noted James Gosling today at Sun's World Wide Education & Research Conference in New York City when asked if Java was in any kind of danger from the newcomers. 'PHP and Ruby are perfectly fine systems,' he continued, 'but they are scripting languages and get their power through specialization: they just generate web pages. But none of them attempt any serious breadth in the application domain and they both have really serious scaling and performance problems.'


As a language, PHP and Ruby may not have the breadth of Java on a single machine - after all, they just 'generate' web pages. Well, they do a little more than that I suppose, they can handle incoming Web requests too. And what applications can you build and at what scale if all you can do is process messages over a network? I suppose the breadth of the applications possible is the same as the breadth of the network your messages travel over - which in this case is global. Maybe someone should tell him that 'the network is the computer'.

March 08, 2006

Alex Russell on Comet

From Phil Windley - this a
very good summary of the web-native event notification system that has evolved from Ajax - Comet. It hits all the points that KnowNow dealt with five or six years ago.

One nitpicking point is this "The problem is that each connection takes a process or thread and there might be thousands of them. Comet can reduce load but not on your current Web infrastructure." A server doesn't necessarily have to allocate a thread or process for each connection - if you look at IRC servers, they handle many thousands of concurrent connections. What is needed is non-blocking IO and Phil's post later points to SEDA - staged event driven architecture. Which brings this full circle from my work at KnowNow that used persistent connections to the browser, to my team at Amazon that used SEDA in a new real-time publishing system we built over the past few months.

It's great to see everybody up to speed on internet-scale event notifications - the power of event driven browser applications is now clear, XML messaging via Atom/RSS is widespread, even Microsoft has contributed algorithms for supporting messages criss crossing a cross-linked web of subscriptions. What's next? I think the 2-connection limit needs to be solved and that will be the final step in the next wave of live Web apps.

March 07, 2006

Simple Sharing Extensions

The Simple Sharing Extensions built by Ray Ozzie at MS is
a pretty decent description of nearly Web-native pub-sub message handling. Anyone doing event notifications will recognize most of the issues and algorithms - there are some very useful wrinkles when considering the context of RSS and feeds.

One thing that this doesn't improve on is the retrieval of data from a single URI - multiple consumers have to really keep up their polling. An alternative is to use something similar to Atom's next/prev linking (which I'm not sure is in the final Atom spec, but it's in the discussion history somewhere). In the case of MS's SSE, if each incremental feed was located at a unique URI, and that feed contained a link to the next set of changes (that the server chooses to hand out) then a lot more caching can take place and the server can choose it's own approach to throttling clients.

March 06, 2006

Comet - event driven Ajax arrives

This
Ajaxian post introduces the name 'Comet' to describe an event-driven approach which is a natural evolution of Ajax applications. It's been great to see so many applications use Ajax and the recent appearance of 'unsolicited message delivery' to browsers is also very cool. There have been a few groups doing this for many years, and it's exciting to see it all about to really take off and become fully mainstream.

The remaining technical hurdle for this technology is the client-side 2 connection limit - at KnowNow we used a 'journal' topic to funnel all messages through, but I don't know what others are doing. I'm going to guess that browsers will soon move this to 4 or more connections. It may be necessary for some deep thinkers to analyze the TCP/IP issues and address it at the protocol level, but I hope not.

Just the other week, I built a simple servlet at work that delivers messages into a browser using a persistent connection. I hooked the servlet to our recruiting/resume pipeline, so now my team gets live notifications the moment a candidate is entered into the system - we now beat out other groups in tagging the candidate and get first choice of many more people. A simple tool, but super useful and it only took a few hours of work, compared to a few weeks of work back at KnowNow so many years ago. Having a built-in XML parser on the client and a robust browser (Firefox in my case) takes care of most of the client infrastructure.

The thing I like about my simple messaging server and client is that it uses hyperlinks identify the set of messages and the client simply 'follows the links' to get more messages. When a client gets to the end of the list of messages, the servlet holds onto that request until more messages arrive - at which point they flow down to the client, achieving fairly low latency. The implementation isn't scalable or efficient, but the approach of using hyperlinks is.

Here's an example chat app (don't expect greatness...)

February 24, 2006

Eliot Blogs!

Finally - Eliot has a blog! Dr. Macro's XML Rants

I worked with Eliot several years ago (at DataChannel) and enjoyed every minute of it.

February 17, 2006

More SOAP apologetics

In "Don Box on Pragmatism and Web Services", Dare quotes Don being pragmatic about when and where to SOAP
Some of the decisions [...] are simple business decisions that are determined by who your target audience is.
1. If you want a great experience for .NET/Java devs, you’ll typically publish schemas (through wsdl) and support SOAP.


Hmm, "business", "audience" and then "developer" just don't flow for me. The market for true Web applications are users, not developers. Developers using early binding procedural languages will be a shrinking market - not a growing market. Those who build for themselves are the new 'developers'.

February 09, 2006

Listings from the edge

This article by Rob Hof about Edgeio is interesting - the vision of decentralized listings is a strong one, and very exciting.

Although Teare's demo was on blogs, that doesn't appear to be the only target. He says this will work for single items on blogs all the way to millions of items from Web stores like Amazon. He contends that RSS feeds reduce the usefulness of centralized repositories of information--most big Web sites today, in other words. "EBays and Craigslists become unnecessary and the tolls they charge become unreasonable," he says.


Well, it looks like Edgeio is a centralized system, but built from de-centralized data that is also directly available. Sort of like syndicated listings.

February 07, 2006

100songs

I've started another blog - not much related to Amazon or Web technology - called
100songs. Since I have a mini ipod and buying music online is easy (and getting easier month by month) and I had a $100 gift certificate from Christmas, I thought I'd buy one song a week and write about it. Feel free to drop by and make suggestions.

January 24, 2006

Windows Live beta agreement - cost of shipping?

I'm checking out the Windows Live email beta, and their license agreement is of course scary, but this really was intriguing:

3. COST OF TESTING. There is no charge to Recipient for testing of the Product. Microsoft shall bear all direct freight expenses relating to the shipment of the Product to Recipient’s place of business and Recipient will pay any return freight expenses.


They unfortunately don't specify how I can initiate the shipping of Hotmail to my place of business - and I'm not sure Amazon would appreciate it even if they did.

January 09, 2006

France July 2005


France July 2005 324
Originally uploaded by dierken.

Last year we traveled to Europe and spent a week in the South of France. July was hot, dry and we loved it there. This is a photo just outside the apartment we were staying in. It's on the edge of the village of Le Garde Freinet.