August 02, 2009

Faster MySQL B-Tree performance

In our ongoing quest for low-cost, high-performance solutions for our platform I've run across TokuDB from TokuTek. This is a MySQL storage engine that uses a different indexing technology that makes updates to an index faster. Making index maintainence faster means being able to have more indices and that allows for richer queries and more interesting applications and analytics.

For one part of our system, I'm interested simply in faster inserts to the index. Having more indices for each table would help in other areas but I haven't started working on those yet. The table in question manages a many-to-one mapping of tokens (text) to internal database row IDs (integers). Essentially, it's a simple key/value lookup table. Currently a single machine does 150 inserts per second, which is fairly abysmal for a table with only 200M records.
This system uses MySQL and the InnoDB storage engine which uses a B-Tree indexing approach. Since the lookups are simply key/value lookups there really is no need for a B-Tree - a hash indexing scheme would work, but unfortunately InnoDB does not support that. We may migrate to Tokyo Tyrant and Tokyo Cabinet since that supports hash indexing and also seems to have better concurrency support (many non-blocking concurrent requests). The keys inserted are almost sequential, but from what I've read the sequential nature of the inserts should help (although the bottom-up building of a B-tree could be optimized by detecting that the key causing a page split is greater than all the values in the page being split - I think Oracle does this).

While looking for hash based indexing for MySQL I found TokuDB which uses the sexy phrase 'fractal tree' for their indexing. Their site has a few introductory whitepapers but a for deeper understanding I wanted to get to the theoretical foundation of their technology.

I soon found that there are several meanings to 'fractal tree' data indexing
  • fpb+ tree. Fractal pre-fetching b+ trees, embeds cache-optimized trees within disk-optimized trees. Unrelated to TokuDB
  • Fractal Trees (TM) from TokuTek. Uses 'cache oblivious' algorithms to improve B-tree usage.
One of their founders was kind enough to post some links on the mostly useful MySQL performance blog - here are the links (however, a few seem to be missing)
The core algorithms seem to be for 'cache oblivious' b-tree structures, based on a description from a masters thesis from 1999. Maybe not all things 20th century suck (postscript files excluded).
Although I haven't finished reading all the papers, what I like about this approach is the way they model algorithm performance as the cost to transfer blocks of data between layers of storage and also that they consider multiple layers. The most important transfer costs being between disk and memory, but including considerations like an OS managed disk cache is good. This models the multi-layer caching that nearly all large scale systems use. We are big fans of measuring algorithm performance and selecting the right tool for the right job.

As an aside, one of the comments on the MySQL performance blog pointed me to InfoBright, a column-oriented store, which might be useful in some of our analytics systems for ad-hoc reporting.

July 26, 2009

Cannon Beach

Last week we spent a great weekend down at Cannon Beach, Oregon. Although there was ocean mist in the mornings, that burned off fairly quickly. The sunsets were fantastic and the kids (and our niece and nephew) had a blast.

See and download the full gallery on posterous

Posted via email from Kinetic

July 17, 2009

Hiking to Lake 22

It has a funny name but good old Lake 22 has a great trail with giant old growth cedars following a stream up to a beautiful alpine lake.
This is a fairly popular hike so there were more people than I'd like but it was good to see so many people with kids and dogs out and enjoying the mountains.

See and download the full gallery on posterous

July 08, 2009

More rock and roll

These guys kick ass compared to the previous band. Love the Boy Scout shirt on the lead singer. The guy on the right had a '59 Les Paul goldtop. Wicked.

Turns out, their name is 'Join the Fight'.


Posted via email from Kinetic

June 25, 2009

Butterfly in the garden

I saw this butterfly flittering between the flowers and ran over to capture it with the camera. It was really big, bigger than I had thought was normal - I wonder what kind of caterpiller turns into this kind of butterfly.

Posted via email from Kinetic

June 17, 2009

More rock and roll at Dakota

Another great night at the Dakota.
This band is Ambrose - any band that starts with a cowbell is okay by me.

Posted via email from Kinetic

June 06, 2009

Fishing

Fishing at the north end of Lake Washington. Nothing here but nice to be outside.

Posted via email from Kinetic

May 30, 2009

Wasting time in the sun

Ready to lay in the hot hot sun and listen to the birds in the forest. probably will fall asleep and burn the other side of my legs.

Posted via email from Kinetic

May 24, 2009

Walking at Juanita Bay park

We haven't been to Juanita Bay park in a long time and after a nice sushi dinner in Kirkland for Rinneke's birthday we headed down there.
It's just as picturesque as I remembered, but it had changed a bit -a few less trees, the beaver dam is gone and I didn't see any turtles this time. Still a log of red-winged blackbirds though.

Posted via email from Kinetic

May 23, 2009

Fostering kittens

We picked up four kittens today that need a foster home for the next few days before they have surgery and then can find a permanent home.
One of our other two cats is very curious (but probably just looking for soft kitten food) and the other cat just hisses as it walks by the room with the kittens.
If you would like to foster kittens - and it's pretty easy - just contact Homeward Pets in Woodinville.

Posted via email from Kinetic

May 22, 2009

Real-time Web just around the corner

The ReadWriteWeb blog has a good post about gathering momentum for a resurgence of interest in real-time search and notifications. I don't think the examples he points to will push it into the mainstream - that functionality has been around in many forms for many years (I even built searchalert.net seven or eight years ago to do that). I do think something will happen, but I'm not sure what application of this technology will make it to the big time.


The Real Time Web is coming so fast we've hardly had any time to think about it yet. So let's do that, shall we? The two hottest technologies online, Twitter and Facebook, are fast integrating real-time delivery of activity streams to their users. Paul Buchheit, the man who built the first versions of both Gmail and Adsense, says the real time web is going to be the next big thing. Buchheit's FriendFeed is a key point of innovation in real time. Social media ping server Gnip promised to turn everything online into Instant Messaging-style XMPP feeds, and though that's been put on hold in favor of more immediately clear value - we've still got our fingers crossed.

May 20, 2009

Beers at Fado

Having a quiet time at Fados and having a Fat Tire.

Posted via email from Kinetic

The end of the road

Out taking the dog-in-law for a walk and liking how everything is that glowing spring green color I love so much. It's like there is this whole new beginning just around the corner

Posted via email from Kinetic

May 18, 2009

Another great weekend at Orkila

We had another great weekend at Camp Orkila on Orcas Island. The weather
couldn't have been better and the sunsets were beautiful as always. The best
part was being able to go out on the water in kayaks - something we've
wanted to do for years. We'll be back next year but that will likely be our
last year with the Y-Westerners program.

Posted via email from Mike's posterous

May 17, 2009

Having a campfire on the beach

Campfires on the beach are one of my favorite things about our campouts. Our local band of hooligans enjoyed the smores and chased each other til it was too dark to see.

Posted via email from Mike's posterous

Cliff climbing on Freeman Island

On the way back from our kayak trip we stopped at Freeman Island. There is a geocache at the top of the cliff. But there was a nest of bees too, so I didn't do the climb this year.

Posted via email from Mike's posterous

Kayaks ready to go in the Sound

After six years of camp on Orcas we finally had the chance to take out some kayaks for a quick afternoon trip.
There were seals sunning on rocks at Doughty Point but no other wildlife.
We were soaked by a freak tsunami caused by a quake that measured 2 on the (Ron) Richter scale).

Posted via email from Mike's posterous

May 16, 2009

Life and Death in the Forest

A game where some are herbivores, some are omnivores and some are carnivores.
You make your choice then run like crazy.

Posted via email from Mike's posterous

May 15, 2009

On our way to Orcas Island

This weekend Stephan and I will be going to Camp Orkila for our annual Y-Guides campout. Hopefully we'll be doing kayaks this time.

Posted via email from Mike's posterous

April 23, 2009

Above the Clouds whitepaper

Here's a whitepaper on Cloud Computing from the UC Berkeley RAD Lab. - just what everyone has been waiting for, a whitepaper on Cloud Computing.

In part this describes obstacles and opportunities. My personal favorite :

  • Obstacle #6 : Scalable Storage
  • Opportunity #6 : Invent Scalable Store


  • That's right, we finally have the go-ahead to Invent Scalable Store.

    The paper gets better the more you read. Another great quote:

    Google Search is effectively the dial tone of the Internet: if people went to Google for search and it wasn’t available, they would think the Internet was down

    March 26, 2009

    When not showing an ad is better

    The Mike on Ads blog had an interesting post a while back referencing some research by Yahoo about how not showing ads might be better for you.


    Instead of showing crappy CPA offers the publisher should show either nothing at all, or some relevant site content. Show a snippet of the friend-feed, or maybe a list of 'online friends'. Show "interesting related links", or "new photos posted"… it doesn’t really matter. Show something that is of interest to the user. The point of the exercise is to train the user to start looking at this specific space again.

    [...]

    If this is obviously so good, why is nobody doing it? Well there’s only one small insignificant problem… Publishers have no way of identifying the top 20% of impressions. You see, especially on social networking sites a huge portion of that 20% are impressions that are sold behaviorally via ad-networks and exchanges. Those ad-networks and exchanges need to see the full 100% to be able to cherry-pick the 10% that are valuable to them thereby making it quite difficult for the publisher to “not show ads” on worthless impressions.


    Showing something other than ads when there is no money involved is a great idea. Unfortunately, most traditional ad networks have no interest or capability to do this. Even 'behavioral' targeting folks aren't in a position because they have only a few 'behaviors' rather than a full tagcloud of interests.

    Our Others Online affinity profiling system has behaviors, interests and a measure of the commercial value of those interests which means we can power ad units that know when its the right time to show an ad, or whether it's better to show relevant news or other content.

    March 17, 2009

    A clean desktop. The hard way.

    My hard drive has been complaining recently and yesterday gave up the ghost. I found a few decent replacement drives online, but needed a new one right now. We had a gift certificate to Staples and what do you know - they had a reasonable 500GB SATA drive for $80. I called them last night just before 9pm - not only were they helpful late at night, but they put one behind the counter for me. Nice people there.

    This morning I backed up all my files to my Linux server here at home and swapped in the new drive. I re-installed Win XP Pro and my machine was now a blank, default machine with no drivers for all the hardware plugged into it - like the big monitor, the speakers, power management, etc. I finally figured out Dell's horrible user interface for installing drivers and soon was back to a running machine with a clean desktop (first time in several years).

    The only oddity was the clock was off by an hour. The old Win XP install didn't know about Congress changing when Daylight Savings Time started.

    After restoring all my files, I needed to update Win XP and install all the old applications again. Not too hard, but time consuming.

    Here's the order that I restored things

    • Dell drivers
    • Firefox browser
    • iTunes (so I could have music while restoring everything else)
    • PuTTY
    • PasswordSafe
    • Windows XP service pack 3
    • Microsoft Office 2003
    • Java 6 SDK
    • Eclipse
    • Apache 2.2
    • Tomcat 6
    • DBVisualizer
    • TortoiseSVN shell
    • GTalk

    March 02, 2009

    Yahoo Query Language and Open Tables

    I've been looking at the Yahoo Open Data Tables and Query Language documentation. This is truly amazing stuff! It provides a service API that accesses many well known data sources (many are Yahoo) and transforms the data into XML or JSON. The data sources can be external URLs that provide XML and Yahoo does the fetch, parse, extract and transform that you want. You can provide a definition of some other external data source and they will hook it into their unified API fetch/query/transform service.

    Some of the data sources are Flickr, local listings, geo location info, web search, image search, news search, weather and so on. One stop shopping for lots of great data.

    Their console http://developer.yahoo.com/yql/console/ is a great way to see what's possible.

    This is what I've wanted for many years. A long time ago I wanted to build a service that would provide "XML data sources" (I even registered xmldatasource.com) for everything available on the Web - now it looks like Yahoo has actually done it.

    Let's hope they keep this data access service open to all.

    February 19, 2009

    Mining the Twitter Stream

    From the Data Mining: Text Mining, Visualization and Social Media blog -
    The Business of Mining the Twitter Stream

    Good post about mining the Twitter stream which hints that existing social media mining companies may already be too established to be replaced by newcomers. This is fast moving area that almost looks like a microcosm of the 'innovators dilemma' - except there is no large, fossilized ecosystem in place.

    SDB adds aggregate function

    Looks like Amazon has added an aggregate function to SDB - basically count(*).

    Amazon Web Services Blog: New Count Function

    I wonder why there isn't a mass quantity of blogs with performance data for various queries against various types and sizes of data sets.

    February 10, 2009

    CouchDB: Jeremy Zawodny's impressions

    I haven't done any work with CouchDB other than read through documents, so I hope Jeremy continues to post what he learns.

    Playing With CouchDB: First Impressions (by Jeremy Zawodny)

    January 29, 2009

    Scalable, reliable key-value lookup service

    Here's a great summary of most of the available key-value storage services from Richard Jones of last.FM - great stuff.

    From what I can tell, Scalaris is only memory-resident at the moment and doesn’t persist data to disk. This makes it entirely impractical to actually run a service like Wikipedia on Scalaris for real - but it sounds like they tackled the hard problems first, and persisting to disk should be a walk in the park after you rolled your own version of Chord and made Paxos your bitch.


    I'm leaning towards Voldemort, but need to look into this more.

    January 05, 2009

    McSweeney's : Fire: The Next Sharp Stick?

    Good gravy, this gave me a great laugh!

    Fire: The Next Sharp Stick?

    You've got to read the whole thing, but here's an excerpt


    ONE: Hairy One, Maker of Fire. Maker of Fire, the Hairy One.


    MAKER: My pleasure, Hairy One. I've followed your work with Ten Men for a long time. It's a remarkable firm.


    HAIRY ONE: So you're the one with the fire?


    MAKER: Yes.


    HAIRY ONE: Is it here?


    MAKER: Well, no.


    HAIRY ONE: Where is it?


    MAKER: Well, in a sense, Hairy One, fire is everywhere. Rather than being an object, say, like your sharp stick, it's really a process, so it can't really be said to exist anywhere. In a sense, fire exists in its own imaginary, virtual space, where we can only talk about what is not fire and what might become fire.


    HAIRY ONE: Whoa whoa whoa! English, please!

    December 31, 2008

    I Hate Computers - Log entry #762

    Here are some helpful tips, if you ever find yourself using 'computers', especially the Windows species.


    • Never ever disable the "VgaSave" video adapter. Ever. This is the fallback software used by Windows to display things on the screen if no other video driver works. If this is disabled, Windows will start but your screen will show nothing but the finest shade of black. If you fail to follow this advice, be prepared to follow these instruction to re-enable vgasave service from the Windows recovery console. Not that it worked for me, but hey, good luck with that.


    • If you perform a 'repair' installation while your screen shows nothing but the black darkness which has enveloped the heart of every poor Windows operating system developer, hoping for your video adapter driver to be repaired, do not (no not ever) turn the power off. Not even once, just for fun.




    • If you happen to boot up your Windows computer and it fails to start due to 'Registry is corrupted' or some such, and you choose to re-install Windows rather than becoming a Tibetan monk (who would likely have fewer problems than a Windows user with a failing disk drive, even considering the Chinese government's approach to freedom), be prepared for a long stretch of file recovery and application re-installation. I recommend Fat Tire Amber Ale from the New Belgium Brewing Company.



    • If you happen to have not followed these tips and you have re-installed Windows and you now have a default user account with a default green and deceptively bucolic grassy field, where before you had many files and folders and possibly (if you are reading this) insipid desktop wallpaper, here is an actual, useful tip : you can change the settings for this newly created user account to use the folders and settings of your previous user account. This won't solve all your problems, but really, do you think that's even possible? Even with advice from me (and I'm composed of nearly 100% pure awesome). Here is what you do :

      • use the regedit program (and if you don't know what that is, give up now) to view HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows NT\CurrentVersion\ProfileList

      • Look for a sub-entry that has a really long value with an entry of ProfileImagePath that points to the fairly empty and useless 'new user account' (e.g. "%SystemDrive%\Documents and Settings\Mike.NAUTILUS". Change the value to point to the location of the old and wondrous user account (e.g. "%SystemDrive%\Documents and Settings\Mike").

      • Go back to the ProfileList registry key, and update the "AllUsersProfile" and "DefaultUserProfile" settings - you'll probably want to continue to use the old folder that had settings for all the old applications you had previously installed and which are most likely useless, since you've re-installed Windows.

      • Log out, then log back in. Hopefully you now see your old desktop. In either case, remember - Fat Tire Amber Ale from the New Belgium Brewing Company.

      • To be honest, you could just copy the files from the old user account folder into the folder for the new user account. But I'm lazy. Just be aware that by pointing the new user account to the old user folder, there may be permission issues with accessing some of the folders or files. I haven't seen that happen but it seems a likely next failure.



    December 09, 2008

    Catalina - point-of-sale ad network

    I wonder if they support cookies...


    Armed with two years of purchase data for 80 million individual consumers, Catalina Marketing is this week launching a new in-store ad network called the Pointer Media Network.

    The information comes from frequent shopper cards covering most of the nation's supermarket chains, thousands of drugstores and other retailers.

    [...]

    Pointer Media works like this: Catalina has installed color printers at the checkout counters of close to 50,000 stores around the country that are linked to the company's massive database of consumer purchases. When a shopper's order is rung up, the printer instantly creates a print ad on a receipt-size piece of paper based on the unique purchasing history of that shopper. The ad is handed to the shopper along with the receipt for the current purchase.

    September 28, 2008

    SpaceX reaches orbit

    SpaceX became the first private company to launch a liquid fueled rocket into Earth orbit. Totally awesome! This was their fourth launch of the Falcon 1 configuration and was carrying a test payload. Their third launch two months ago didn't achieve orbit and actually was carrying paying payload - oops.

    This is a huge step on the road to space.

    August 02, 2008

    SpaceX - Flight 3 of the Falcon 1

    SpaceX is launching their Falcon 1 today from Kwajalein - about 1 hour to go before liftoff, if things go according to plan.

    Here's a live webcast - http://www.spacex.com/webcast.php

    July 30, 2008

    WebHooks

    This looks interesting - in a 'teach people how the Web really works' kind of way. WebHooks is a catch phrase for Web application development where notification are sent from theh source to the listener via HTTP POST, rather than the other way around via polling (which, as some have said, doesn't scale).

    Somewhat related to my earlier post on how HTTP can be used as an alternate to XMPP/Jabber in a publish/subscribe scenario without too much problem or angst.

    July 29, 2008

    MySQL and LAST_INSERT_ID()

    We've recently migrated our software servers from a dedicated server environment (BlueGecko - really good people and service) to Amazon's EC2 'virtual compute cloud' environment.

    The new system has a very nice performance monitoring capability based on Ganglia that gives us visibility into the performance of service requests as well as details on more fine-grained functions that take place within each request. We can now see not only the number of requests or functions but also the average duration and the time it took for 50%, 95% or 99% of the requests to complete in a five minute interval. This percentile breakdown gives a quick feel for how 'spiky' performance is and how common outliers are.

    So far, things have gone very well but there was one function of the system that seemed to suddenly have terrible performance. I was able to quickly see which area of the code was involved, which pointed me to some SQL statements.

    As our system adds new anonymous profiles, the code retrieves the unique number assigned by MySQL using the special query that MySQL provides:
    SELECT LAST_INSERT_ID() AS user_id FROM oo_anon_user


    However, it turns out that this SQL is actually incorrect - the extra "FROM oo_anon_user" caused the database to return every single record from the table back to the client, on every request that created a new record. This of course took some time.

    The correct syntax to retrieve an auto-increment field from a MySQL table is
    SELECT LAST_INSERT_ID() AS user_id
    Remember to omit any FROM clause.

    It's much much much faster now.

    July 24, 2008

    REST and pub/sub

    It's unfortunate that technologists continue to propagate serious mistakes like "[...] its also clear that REST and its inherent polling mechanism isn't the best way of building a user notification system [...]"

    REST is about state transfer - and event notifications are also state transfer.
    As for HTTP, it isn't only "polling" - anyone that has posted a blog entry knows that. The 'client' can 'post' updates to the 'server' - exactly the same as event notifications via XMPP. The great thing about XMPP is the federated multi-hop capability with 'trust' built-in. Just like email, only with everyone using settings for very low latency delivery.

    There have been multiple publish/subscribe over HTTP mechanism (comet, mod_pubsub, KnowNow, etc) over the years.

    July 21, 2008

    Twhirl Adds Identi.ca

    I don't really follow the social/friend feed streaming application space, but this post about Twhirl from
    ReadWriteWeb had a quote that caught my eye (emphasis mine):

    However, for the regular user, always on social networking doesn't have to be a source of stress - it just means that when you go online, socializing with others just becomes part of the overall experience of being on the internet.

    June 01, 2008

    MoonBoy

    We've been doing GeoCaching for a number of years, enjoying learning about new parks and finding Geocaches on our trips. Earlier this year we created our own Geocache near our house and placed a Travel Bug in it - MoonBoy. His goal was to view a space shuttle launch and to someday travel on a Space Shuttle.

    Amazingly, he's made his first goal and was an observer to this weekend's launch of STS. This is totally cool! Here are the photos on the GeoCache site

    May 12, 2008

    Oh, the irony of shallow WSDL

    I'm looking into the API for the Hi5 social network and unfortunately found some WSDL. They also have a REST API, but it's documentation appears auto-generated from WSDL that nobody actually filled in. Somewhat useless.

    Normally I wouldn't post about WSDL, but I couldn't pass up the irony of the WSDL documentation for the authentication service API having an HTML form describing how to authenticate the user. If that's not irony, I don't know what is.

    May 09, 2008

    Ultimate Twitter revenue model - chatbots??

    From ReadWrite Web

    "Essentially, this would entail Twitter parsing over the Tweets of a given user, as well as the Tweets of the users he/she is following. Common keywords, themes, and phrases are then pulled from this data and associated with that user. As a result, highly-targeted ads can be displayed based on the user's network of content ("web design", for example). These simple text ads would look very similar to regular Tweets, but would be clearly marked as "Sponsored Content"."



    I think chatbots haven't work for a reason - people want to chat not shop.

    Reading RWW and other pundit blogs that describe "how the future will work" reminds of reading Popular Science as a kid and gazing in wonder at the flying cars and transparent house soon to be built.

    May 06, 2008

    Web scale pubsub

    Looks like people are now starting to talk about a de-centralized, Web native pubsub system. Now is the time, if applications are going to avoid being based on a single company's service. Too bad pubsub.com and KnowNow aren't driving this.

    April 24, 2008

    CouchDB

    I'd heard the CouchDB name but but didn't realize that it's an HTTP-accessed, Erlang-implemented, schema-free, distributed datatabase - and it supports Javascript for defining views. Time to dive in!

    From Apache CouchDB: The CouchDB Project:

    "Apache CouchDB is a distributed, fault-tolerant and schema-free document-oriented database accessible via a RESTful HTTP/JSON API. Among other features, it provides robust, incremental replication with bi-directional conflict detection and resolution, and is queryable and indexable using a table-oriented view engine with JavaScript acting as the default view definition language."

    April 21, 2008

    What programmers like

    Quote of the day, found on the High Scalability blog while looking into Amazon's SimpleDB
    "Programmers like problems they can solve with more programming."

    April 20, 2008

    Online services and availability

    I've been working with the design and development of online services for quite some time and am very familiar with what 'high availability' means, but today I had the opportunity to learn first hand how the lack of reliability can directly impact people. This is a minor example, to be sure but as more 'software as a service' companies are built out more people will be affected in ways that they don't quite understand.

    My daughter attends a Middle School that uses the online curriculum from Key Curriculum Press - I have no idea if the content is decent or not (but knowing the teachers at the school I'd bet it's good) and in order to access specific materials a login process has to take place and then access to some PDFs become available. This weekend, the login service was unavailable and so my daughter could not work on that part of her homework - very annoying.

    The problem is likely something as simple as a server misconfiguration or a host being down (I can't imagine there's too much horsepower required for this sort of service), but as a 'customer' we are unable to do anything. There is no 24x7 tech support or online contact information. I could make a phone call, but that would only work between 8am and 5pm (PST), and I'm pretty sure kids do homework outside of those hours.

    I've sent an email to their president and we'll see what kind of response that gets.

    Update - I received a nice email from their president and she also sent a follow up email with more details. Very nice!

    April 19, 2008

    MyOpenID for Your Domain

    This post from ReadWriteWeb about
    MyOpenID for Your Domain is interesting. I've looked into OpenID several times in the past but it didn't have enough immediate value to get me to do anything about it. Now that more sites are 'supporting' OpenID and with the simple to use MyOpenID provider maybe I'll set something up. It would be interesting to see decentralized profiles and authentication, it seems a shame that everyone is getting locked into a few social networks for hosting and managing personal profiles and shared contact info.

    April 16, 2008

    Yaws and reverse proxy

    It seems that there are a lot of situations where efficient connection handling in an HTTP reverse proxy is needed to help with scaling and keeping costs down, and my first guess was that Erlang would be the way to go.

    So I just searched to see if there was a reverse proxy written in Erlang and sure enough Yaws has one.

    Next, I need to find some folks that have actually used it and ask some questions. Especially for persistent connections (accepting a request but sending no response until there is data available, like in a message queue or pubsub scenario).

    April 14, 2008

    Of Liquidity, Competition and Platforms

    This post by Bob Wyman "Of Liquidity, Competition and Platforms" reminds me why I subscribe to tech blogs - occasionally you get a deep thinker that clearly and concisely describes something that likely took them a long time to figure out.

    "As many others have said, much of what people are building today is 'features' not products. As long as that is the case, raw economics is the real problem with the software business. Competition strengthens the platform builders while eviscerating the component builders... The rich get richer and the builders of innovative features must be satisfied with the 'personal rewards' of doing a job well unless they get lucky in the buy-out lottery game."


    With the various cloud computing efforts well underway, and with each providing a different approach as a 'platform', it would be good for developers and entrepreneurs to think about which road to travel.

    April 10, 2008

    More REST and Erlang

    Man, I wish I had the time to dig into these posts about Erlang, Yaws and RESTful services.

    Cross domain web apps

    This post about cross domain Web apps is spot on. The growth of the current Web was due to millions of silos - no data shared between sites. Over the past few years technology like iframes, ajax and other styles of 'mashups' have made it possible to create applications composed from cross-domain services. I believe this will drive the next major growth curve of the Web.

    April 04, 2008

    Elephant on a Bicycle



    Here's an interesting blog on China from someone living there. The photos are both beautiful and intriguing.

    Elephant on a Bicycle

    An example of the writing (which I really like)

    I'm now in Xining, Qinghai, in a youth hostel 15 floors up - which you think would offer remarkable views of the mountains that crowd the city on all sides. But it doesn't, this is China. 15 floors up generally just presents you with worse pollution, as it seems to procrastinate in mid-air limbo.

    I'm in a quiet little room with two computers. Beside me sits a young monk of about 18, who has rather rashly decided to accessorise his austere crimson robes with a pink feather boa. I couldn't make that up. Furthermore, he is currently downloading provocative pictures of Nicole Kidman; certainly harder to find fault with this inevitable phase of teenage experimentation.

    Xining city is a fascinating stopover because it offers a unique diversity. It is still predominantly Han Chinese, in both it's populace and it's blindfolded embrace of modernity - noisy, dirty, difficult to look at, pulsing
    with energy. But alongside this, it contains a deeply religious aspect from the presence of size-able pockets of Islamic and Tibetan communities. In places there even exists a meditative quiet - a stillness - which amidst the commercial bustle and the perpetual foot-race of human progress can seem almost revolutionary. This is especially apparent in the areas surrounding the great mosque where I spent most of the day after befriending a young Koranic scholar named Farooq.

    March 21, 2008

    Party like it's 1994

    You learn something new every day. Although sometimes it takes 5,110 days.
    From Steve Vinoski's blog - here's
    A Note on Distributed Computing from Sun in 1994.

    We argue that objects that interact in a distributed system need to be dealt with in ways that are intrinsically different from objects that interact in a single address space. These differences are required because distributed systems require that the programmer be aware of latency, have a different model of memory access, and take into account issues of concurrency and partial failure.

    March 11, 2008

    What Sucks About Erlang

    Very funny yet enlightening post about Erlang. And I still want to build a real system with it.

    From Damien Katz: What Sucks About Erlang:


    "The only purpose of the true -> ok line is to give it an else condition to match. That weird taste in the back of your throat? It's probably vomit."

    March 01, 2008

    Seattle's nuclear reactor


    I knew that some universities have or had nuclear reactors for research, but I didn't know they could look this beautiful.


    From Crosscut:
    Abby Martin hopes that, whatever its fate, the Nuclear Reactor Building finally gets public acknowledgment of its role in history, and credit for its architectural originality. That would be in keeping with its original intent. It's hot core may be gone, but it can still teach us something.

    February 25, 2008

    No more ugly desktop software

    From
    ReadWriteWeb it seems Mr Kirkpatrick is all a quiver with Adobe AIR :
    AIR lets developers use Adobe Flash, Adobe Flex, HTML and AJAX to create desktop apps. That means no more ugly desktop software!


    No more ugly desktop software. Do you really think it was the software that was the problem before? Unless you're part of the short-attention-span-journalism industry, it'll take more than drop shadows and bouncy icons to get past ugly software.

    February 22, 2008

    Dare Obasanjo aka Carnage4Life - How "View Source" Broke the Web

    I haven't been following the Web blogging elite recently (really, 'view source' broke the Web??), but this post - How "View Source" Broke the Web - made me think of Bob Dylan.


    Broken lines, broken strings,
    Broken threads, broken springs,
    Broken idols, broken heads,
    People sleeping in broken beds.
    Ain't no use jiving
    Ain't no use joking
    Everything is broken.

    February 02, 2008

    SpaceX signs more business, builds more engines

    It looks like SpaceX is continuing to make good progress, both with engineering and with business.
    From the SpaceX status page:

    "For the last few months, optional activities such as website updates have gone on the back burner while we finished the regeneratively cooled Merlin 1C engine, got the Falcon 9 first stage integrated, proof tested and fired, signed up our first GTO (geostationary transfer orbit) customer - Avanti Communications Group, and took our COTS system past the big CDR milestone."


    Signing a contract for placement into a geostationary orbit is huge! I'm really hoping these folks don't spend all their capital on development of bigger engines without establishing the operational business that will generate revenue. They seem to be looking to raise $100M in the first half of 2008, hopefully that will happen and keep this company going to profitability.

    February 01, 2008

    Yowza! MicroHoo in the future?

    Looks like Microsoft is serious about buying advertising's future - they have offered $44B for Yahoo.
    Favorite punditry (from Stowe Boyd)

    "Personally, I think the Microsoft and Yahoo matchup is like two tired swimmers who bump into each other and then wind up drowning each other in their scramble to survive. But Yahoo will be the first to go under in this embrace."


    This one from iMe on Twitter is good too:


    "Microsoft and Yahoo! Its like a blind man trying to lead a deaf guide dog...."

    January 19, 2008

    Stonebraker on MapReduce

    I've too busy lately to post on all the exciting things happening in the database world - more people getting into column-oriented storage, Amazon's SimpleDB service, consumer database Web services like blist and LongJump, Sun buying MySQL. But I couldn't pass up commenting on Joe Gregorio's post on Stonebraker on MapReduce.
    It seems everybody is panning Stonebraker's evaluation of MapReduce as incorrectly comparing it to a DBMS. One commenter even said (ironic comment of the year) Michael Stonebraker should learn what a DMBMS is.

    The point, which I'm sure someone has pointed out, can be found by looking at the summary of their findings. Here are a few:
    - A sub-optimal implementation, in that it uses brute force instead of indexing
    - Missing most of the features that are routinely included in current DBMS
    - Incompatible with all of the tools DBMS users have come to depend on

    These read like a quote from The Innovator's Dilemma. People enjoy the benefits of MapReduce for large scale data processing because of these points. The lack of support for these DBMS features and tools are the reason it scales like a mother fucker. That is the feature people want. And it does it 10x better and cheaper than anything else. Those other things simply don't matter to them. It's a new audience and a new market.

    December 09, 2007

    Animation, Pivot and Maya

    Over the past few months I have been volunteering in my son's gradeschool class to teach art and animation using computers. The class has several computers in the room and the school has a good computer lab with lots of equipment, and of course most of the kids have a computer at home - quite a change from when I was growing up.

    Before I started, I did a little research into available programs for art, sound and animation. There aren't a lot of solid applications freely available, but here are the ones that I found to be interesting:

    • ArtRage - beautiful digital oil painting

    • Pivot - stick figure animation (very addictive)

    • Scratch - visual programming of sprites from MIT (essentially visual logo)

    • Microsoft Movie Maker - simple track-based video and audio compositing



    I found one sound effects generator but haven't found the perfect sound effects program appropriate for these kids. Yesterday I found Reaper, a full featured audio track and effects program - it looks wicked cool, but complicated for gradeschoolers.

    The kids have amazed me by picking it all up really quickly. Several of the kids already knew about the Pivot cell animation program, and others have already made stop motion animation with Legos, digital cameras and Microsoft Movie Maker. Normally I would have expected the boys to be the ones more into working with the computer, but in this class everybody is very excited and has been diving in. All the kids are trying their hand at different skills - some are more comfortable creating a background painting in ArtRage, but still make a go of animation in Pivot. Others quickly grabbed some background photos off the Web and brought them into Paint.net and faded them out to make a good background.

    The teacher has told me that she has to shoo them out during recess because they'd rather work on their animations than play outside! Makes me feel so proud...

    Right now I have them working on a project in teams of two. To explain why working in teams is a good thing to learn, I told them of time I visited Pixar and learned how they used teams - an 'artist' and an 'engineer'. They got a kick out of my describing walking the halls and looking at pencil sketches of a cowboy and a spaceman - I thought the movie idea was great but didn't know if it would have mass appeal. I still regret not having the gumption to ask for one of those sketches - I was so in awe of meeting Ed Catmull at the time, it was hard to speak! I'm sure I made a fool out of myself...

    Recently I've thought about introducing the class to 3D graphics, but it can be very time consuming to create the models and I haven't found a good animation program for use with pre-existing models. I did recently download the personal edition of Maya, and maybe there's something in there but I'm guessing it'll be too much for the kids.

    December 07, 2007

    Edgeio shutting down

    Looks like Edgeio is shutting down. That's too bad, I liked the idea of supporting decentralized listings of offers. I suppose being a centralized aggregator wasn't the way to go.

    Good quotes from Michael Arrington's (a co-founder) comments:

    # Andrew
    December 7th, 2007 at 12:31 am
    what exactly did you spend 5 million dollars on?

    # Michael Arrington
    December 7th, 2007 at 12:32 am
    Andrew - parties, scotch, hookers, blow. you know, the usual.

    November 19, 2007

    The Old Web = Ten Million Catalogs

    Stowe Boyd always has something interesting and thought provoking to share on his /Message blog. If you are interested in the intersection of technology and society, I highly recommend giving a little attention.

    Today's post is
    The Old Web = Ten Million Catalogs (part of his The Social Web: What's The New Web Worth series) which looks back at the past ten years of the Web to draw out the distinction of what the new Social Web has become.

    These services are based on a catalog metaphor, where sellers can offer goods or services, and buyers (generally consumers) can find them and acquire them. The volume and low overhead of online services hollowed out the markets in most areas thet they touched, for example, sideswiping brick-and-mortar bookstores, blowing up the travel agent business, and strongly cratering the head hunter marketplace.


    I like where this is going - pointing out a 'catalog metaphor' for Web services helps me describe to people how my current company isn't just a directory of people - it's a way to add a social dimension to the Web, centered on people not pages.

    November 17, 2007

    Steve Vinoski’s Blog

    Steve Vinoski is one cool dude. I've always had the utmost respect for IONA and even after leaving a while back, Steve continues to show the attitude of professional engineering that garnered that respect. With his measured explanation of his view of REST and Dare's theory and practice post I think I'm ready to walk away from cooling embers of the dying REST .vs. SOAP flame war. Time to unsubscribe from rest-discuss.

    From Steve's blog, here are some good quotes:
    People seem to get really upset when I say that the static typing benefits of popular imperative languages are greatly exaggerated, and when I say that developing real, working systems in dynamic languages is not only possible, it’s preferable.


    Either way, no interface definition language is ever going to keep you or some other real live person from having to figure out what the service actually does and how to actually use it, and then coding your client accordingly.


    Remember, REST is an example of applying well-chosen constraints to achieve desired architectural properties for a broad class of distributed systems, and so that’s what its constraints are all about.


    Either way, anti-REST folks commonly claim that REST’s success is due only to the fact that there’s a human-driven browser in the mix, but that’s one of the dumbest things I’ve ever heard.


    Having a solid thinker like Steve Vinoski blogging makes the Web a better place. Can't wait to hear the details of what he's working on.

    November 15, 2007

    Amazon PR: Neither Open Nor Social

    Looks like there was a bit of a messaging snafu between Marshall Kirkpatrick and Amazon and the results are not pretty - ouch. There's messaging then there's messaging. Let's see what the 'go-forward strategy' is...

    Is this just communication gone awry, or is there a culture clash looming?

    From Read/WriteWeb
    - Amazon PR: Neither Open Nor Social


    Did you know that there's been no RSS feeds for top selling items in categories at Amazon.com? Well, there is now - and they were so excited that they figured it out, that they wrote it up in a press release.

    November 04, 2007

    Nielsen on Generic Commands

    I ran across this interesting bit from Jakob Nielsen's UseIt site about user interface design. It's about using the same few commands in a UI and is an interesting twin to the 'uniform interface' aspect of REST. Although a 'user interface' and a 'network interface' are very different beasts I can't help but wonder if there's some fundamental reason for the similarity.


    Summary: Applications can give users access to a richer feature set by using the same few commands to achieve many related functions.

    In application design, there's a tension between power and simplicity: Users want the ability to get a lot done, but they don't want to take the time to learn lots of complicated features.

    OpenSocial - what? no logo?

    Everybody has been commenting on the news of the week - Google and MySpace spinning out their own system for applets within social networks called OpenSocial.

    It seems everyone has missed the biggest gap in the OpenSocial system - they have no logo! Rather than talk about the disappointing lack of open content in their 'extension' to Atom for profile data (anyone heard of hCard or FOAF?) or the lack of anything innovative like client-side includes (that might actually lead to social networks being spidered by a big search engine) I've decided to contribute my l33t grafx skillz to the community in the form of an OpenSocial logo. It includes the mandatory Web 2.0/startup color scheme of blue and orange. I need to figure out how to include the RSS radio waves into the design somehow...

    or this...

    Next, a motto : "Where do I belong today?"

    October 28, 2007

    Armadillo Aerospace almost wins prize

    So close... Armadillo Aerospace had a couple good flights but tipped over just before touchdown and the next day caught on fire on takeoff. They were this close to winning the Level 1 prize in the Lunar Lander Challenge.


    More about the Lunar Lander Challenge over here.

    October 05, 2007

    Friday humor

    Gotta love inner-thought-humor on a sunny Friday.

    • You want four million users by DECEMBER?? You have four hundred active licenses for your product currently! Nothing - and I mean NOTHING - is going to add four zeros to the end of that number in three months short of hiring Arthur Anderson to handle the bookkeeping.


    • Wait... First you wanted to clone Digg... Then you wanted to "add the social aspects of Facebook to it," and NOW you want it to be Wikipedia? Where the HELL did you spend your morning? In the "Web 2.0 Company Names to Memorize" symposium sponsored by the local Linux Enthusiasts club?


    • Uh... Four million active users means minimum 20,000 concurrent users at any given moment, and you want to do all of this on ONE co-located virtual server in India? On .Net and MS SQL Server? Honestly? You really, really think that's how it will go? In that case, can I punch you? Please? I mean, I only ask because you seem like the type of person who'd ponder the question and then just blurt out "Yes," and I've been dying to hit something since I pressed "1" to join your conference.

    The Web! It's People!

    Tim Bray always has interesting things to say - I've enjoyed the distributed discussion he's started around Erlang - but his latest piece is a direct echo from the Others Online playbook (emphasis added):
    Here’s the thing: the Net’s killer app has always been other people. There are side benefits, like access to all the world’s information. But the links that matter aren’t between pages but people, and they’re strong and rich and subtle. Multiply the infinite flavors in human relationships by a thickening bundle of means-to-connect; that product is what’s new and what’s good and what’s exciting. People who are looking for the Next Big Thing are mostly looking in the wrong places. And anyway, you don’t need to look, it’ll find you.

    October 04, 2007

    Good tips on starting a startup

    This is a very decent (but short) article on things to do and think about if you are starting to develop your own small Web-based software startup. It's from Read/Write Web and part of a larger series. If you have fifteen years of experience, you'll know this already - otherwise take a look.
    How To Create a Web App

    September 27, 2007

    Amazon Payment System

    I was looking at the Amazon Flexible Payment System service API and it's a whopping big set of docs. Unfortunately, I was sorely disappointed to see the operations in the HTTP based API (called a REST Request) seem to all be GET requests with an Action= query term. Ugh.
    The documentation doesn't even mention what HTTP method to use, what content-type to submit or expect as a return, etc. I'm pretty sure that they simply don't know the difference. Sad but true. At least there are docs.

    I suppose they didn't have the time to make it simple.

    September 25, 2007

    NASA Tech Briefs

    A long time ago I used to subscribe to NASA Tech Briefs - a slim monthly magazine full of hard-core engineering articles and light scientific reporting.
    A few months ago a friend had an issue and I decided to see what kind of online presence they have now - and they have a good one. No Atom feed, but still lots of techno-bits!

    NASA Tech Briefs

    September 19, 2007

    Metaplace: open DIY virtual worlds

    This I have to check out... it sounds very much like what I've wanted to build for the past ten years or so.

    Metaplace: open DIY virtual worlds for everyone

    September 17, 2007

    Facebook - the end is near

    Ah, yes. Facebook is on the decline already - when I logged in this week, I got a an ad for Zwinky. Animated, colorful, annoying. Completely not me. Except the annoying part.

    September 03, 2007

    CouchDB: Thinking beyond the RDBMS

    From Assaf at labnotes,
    CouchDB: Thinking beyond the RDBMS

    It stores document in a flat space.

    There are no schemas. But you do store (and retrieve) JSON objects. Cool kids rejoice.

    And all this happens using Real REST (you know, the one with PUT, DELETE and no envelopes to hide stuff), so it doesn’t matter that CouchDB is implemented in Erlang. (In fact, Erlang is a feature)

    Here’s where it gets interesting. There are no indexes. So your first option is knowing the name of the document you want to retrieve. The second is referencing it from another document. And remember, it’s JSON in/JSON out, with REST access all around, so relative URLs and you’re fine.


    Hmm, Erlang. Hmm, no schemas. Hmm, no write consistency. Sounds perfect.

    August 09, 2007

    HTTP errors - a photoset on Flickr

    (From BB)

    HTTP - a protocol so simple, even a cartoonist can understand it. (and why haven't you bought a shirt yet?)

    404 Arrgghh - (insert pirate code joke here)

    404   Arrgggh

    August 06, 2007

    SubAtomic future

    I just had a thought - what if the Atom Publishing Protocol and Subversion had a mashup - we could call it SubAtomic!

    July 17, 2007

    Object serialization raises it's insidious head

    Funny post from Dare Obasanjo decrying the horrible state of affairs created by the company he works for.

    This part is about object serialization and version incompatibility.

    However, the insidious thing is that the failure wasn’t because their application was improperly coded to fail if it saw a fruit it didn’t know, it was because the platform they built on was statically typed. Specifically, the Web Services platform automatically converted the XML to objects by looking at our WSDL file (i.e. the interface definition language which stated up front which types are returned by our service) . So this meant that any time new types were added to our service, our WSDL file would be updated and any application invoking our service which was built on a Web services platform that performed such XML<->object mapping and was statically typed would need to be recompiled. Yes, recompiled.


    Insidious indeed! Who in their right mind could have predicted such things!

    And this...
    It’s sad that as an industry we built a technology on an eXtensible Markup Language (XML) and our first instinct was to make it as inflexible as technology that is two decades old which was never meant to scale to a global network like the World Wide Web.


    Who's this industry of which you speak, Kemosabe?

    July 14, 2007

    Getting all meta

    Dare Obasanjo seemingly always has time to create detailed and thoughtful posts, I've gone from being annoyed at early posts to reading them regularly.
    His latest two posts are worth reading as well, but I can't help but comment on his strawman characterization of ReST.

    I should probably start out by pointing out that the title of this post is a lie. By definition, RESTful protocols can not be truly SQL-like because they depend on Uniform Resource Identifiers (URIs aka URLs) for identifying resources. URIs on the Web are really just URLs and URLs are really just hierarchical paths to a particular resource similar to the paths on your local file system (e.g. /users/mark/bobapples, A:\Temp\car.jpeg). Fundamentally URIs identify a single resource or aset of resources. On the other hand, SQL is primarily about dealing with relational data which meansyou write queries that span multiple tables (i.e. resources). A syntax for addressing single resources (i.e. URLs/URIs) is fundamentally incompatible with a query language that operates over multiple resources. This was one ofthe primary reasons the W3C created XQuery even though we already had XPath.

    That said, being able to perform sorting, filtering, and aggregate operations over a single set of resources via a URI is extremely useful and is a fundamental aspect of the Web today. As Sam Ruby points out in his blog post Etymology, a search results page is fundamentally RESTful even though its URI identifies a query as opposed to a specific resource or set of resources [although you could get meta and say it identifies the set of resources that meet your search criteria].


    I don't think there's any harm in the mis-characterization of ReST and the detail of the post about the two query systems from Google and Microsoft is rather interesting, but I was surprised at his misunderstanding since I thought he knew better. Dare does have an out - as any good program manager does - by saying "... you could get meta...", so I will take that opportunity and get meta.


    • "URIs on the Web are really just URLs and URLs are really just hierarchical paths to a particular resource [...]" - no, they are not really just hierarchical paths. They are an identifier. Sometimes they are a 'path', but usually not.
    • "Fundamentally URIs identify a single resource or a set of resources." - always just a single resource.
    • "A syntax for addressing single resources (i.e. URLs/URIs) is fundamentally incompatible with a query language that operates over multiple resources." - this is core missed direction in the explanation of ReST. The term 'resource' as applied within RDBMS isn't the same as 'resource' as applied within ReST. A ReST 'resource' is the view from the outside, and datbase 'resources' are the data elements on the inside. Read up on Data on the Outside vs. Data on the Inside by Pat Helland (not exactly relevant to this post, but a catchy title).
    • "a search results page is fundamentally RESTful even though its URI identifies a query as opposed to a specific resource or set of resources [although you could get meta and say it identifies the set of resources that meet your search criteria]." - the URI actually does identify the results (as Dare suggests) and not the query. The URI is the query. This means you can operate on the set of data that are the results - you could delete them, retrieve them, replace them, etc.


    I'm not suggesting that URI are better than or a replacement for queries, just that they are different. They do different things and are both very useful, and when you think in terms of "a URI is a query" you think of the processing of the implementation and miss out on some aspects of Web architecture which naturally arise when you think of "a URI identifies the results".

    May 21, 2007

    Ruby on Rails Job Trends


    Job trends for Ruby on Rails From the indeed.com job search site (which looks very nice by the way), the trend for Ruby on Rails mentions in job postings is on a steep growth curve. The absolute percentage is smaller than other terms, and there may be other factors that contribute to the trend but it's pretty telling.


    Job trends for Ruby on Rails and J2EE But to keep things in perspective, compare Ruby on Rails with J2EE job postings



    Job trends for Ajax, Web 2.0, Ruby on Rails More happy fun job trends for Ajax, Web2.0, et al.

    May 20, 2007

    Branding ads and ad listings

    I just saw this post on BoingBoing about this ars technica post on the psychology of banner ads which is interesting in light of Google's and Microsoft's recent acquisition of 'creative/banner ad' networks.

    "The research concludes that repeated exposure to a product via banner ads generates a positive feeling towards that product. The good news for consumers is that a critical reevaluation of the product can make these positive feelings vanish." and
    "This suggests that familiarity-based advertising may work best for impulse buys, where more detailed evaluations aren't likely to occur."

    I found these two papers yesterday when reading up on economic theory and internet advertising. This one Internet Advertising and the Generalized Second-Price Auction - Selling Billions of Dollars Worth of Keywords is the most readable of the two and talks about the bidding process for online ad placement. They cover the history and describe the shift from "pay what you bid" to the current "pay the next lower bid" (called a 'generalized second price auction'). Really very interesting stuff.

    The second paper Brand and Price Advertising in Online Markets looked at different fundamental forms of advertisements - brand advertising and 'price' advertising. I had high hopes for learning from it, but it was very dense with too much lingo specific to this research area for me to fully understand. Here's an example: In contrast to models where loyalty is exogenous, these crosschannel effects lead to a continuum of symmetric equilibria. Yeah, I, uh, was thinking the same thing.

    Their models and assumptions also seem questionable and so I don't know that their results apply to the world we live in, as opposed to the simplified models they used to prove their theory. In any case, the questions being asked and the attempt to find answers are still valuable.

    Here are their findings (and again, these may not really apply to the real world)
    While each firm finds it optimal to advertise its brand in an attempt to “grow” its base of loyal customers, in equilibrium, branding (1) reduces firm profits, (2) increases prices paid by loyals and shoppers, and (3) adversely affects gatekeepers operating price comparison sites. Branding also tightens the range of prices and reduces the value of the price information provided by a comparison site.
    Their research shows that brand advertising allows a firm to have higher prices, since loyal consumers aren't as price sensitive, but their conclusions are that profits are less - which I don't understand, unless the cost of creating the brand is very high. The Ars Technica post shows that other research continues to confirm that viewing brand ads creates a positive impression, which is one step towards converting shoppers into loyal customers.

    This got me thinking about the various forms of advertising. The second paper distinguishes 'brand advertising' from 'informational advertising', which I agree is a useful distinction. I think of it in terms of how actionable the advertising is - how delayed is the payoff. For branding, the payoff is very indirect but could be profitable (or not, depending on which research you subscribe to) due to higher prices or repeat business or cutting out the competition through causing the customer to not engage in comparison shopping. For ad listings that show a very specific product or category, which are often a gateway to a purchasing decision, the payoff is fairly direct. Some even feel that the advertising cost model may evolve into 'cost per action' and go beyond 'cost per click'. However, cost per click is currently much easier to gather metrics in a two-way trusted fashion than measuring the final transaction in a two-way trusted fashion.

    Another example of a very actionable advertisement is the Amazon or EBay offer listings - these are immediately purchasable, and through syndication via Associates are widely used as advertisements. I don't know of any company that does placement optimization of Amazon Associate links (raising placement for offer with better click through rates, better commission for the associate, etc), but I think Amazon has started doing some of that with Omakase links. One interesting thing to consider is that both Amazon and EBay are similar to price comparison sites due to the large number of offers for a single authoritative item (EBay doesn't have very good item authority, but people work around that issue by using manual searches). Getting top placement on an Amazon offer listing page, or in the 'buy box' on the details page, doesn't use an auction bidding process the way Google or Yahoo paid search listings do. The offer listing position is based on the offering price and estimated shipping costs.

    May 19, 2007

    Online Ad industry consolidation

    So, what's up with all the acquisitions of online ad networks within the past 30 days?
    • Google purchased DoubleClick for $3.1B on 4/13/2007 (on revenues of $150M).
    • Yahoo purchased the remaining interest in RightMedia for $680M on 4/29/2007 (on revenues of $70M) They had purchased a 20% stake back in Oct 2006.
    • Microsoft acquired European mobile ad network ScreenTonic on 5/3/2007.
    • Microsoft acquired aQuantive for $6.1B on 5/18/2007 (on revenue of $442M).
    • AOL acquires major interest in Adtech AG on 5/16/2007.
    • AOL acquires mobile ad network Third Screen Media on 5/17/2007.
    • WPP Group acquired 24/7 RealMedia for $650M on 5/17/2007.

    A few billion here, a few billion there, pretty soon you're talking real money.

    What's happening is different for each player, but the overall trend is the same - expanding beyond paid listings into creative branding. Paid search was $6.7B last year and brand advertising was $3.3B. This interest in brand advertising may be a reaction to the expectation that television - which is mostly branding style ads - is moving online.

    In Google's case, they are buying a company that has been successful with creative ads, essentially banner ads. Banner ads are the most common choice for branding rather than being used for actionable ad listings. They also inherit distribution agreements for AOL and MySpace, and more distribution capacity helps draw advertisers into Google to bid for placement.

    Microsoft hasn't done well in any online ad segment - listings or branding - and with their acquisition they will be more involved with holistic ad campaigns and deep in the creative arena of advertising. The includes the ad agencies that do the actual construction of creative ads and create very innovative branding experiences like custom website which are blurring the lines between interaction and advertising. There may be some future tie-in with 'rich internet applications' and Silverlight. I may be wrong, but I don't see aQuantive providing additional distribution capacity - 'inventory' as the ad industry calls it. They of course claim it 'extends their platform'. Everything's a platform to Microsoft. Maybe they should simply try providing value instead.

    The hope is that the contextual and behavioral profiling that is done for ad listings will be applied to better target brand advertisements. A requirement for this to work is for the ad network to know a lot about their audience - something that a single site cannot accomplish. Effective audience profiling is orthogonal to Web sites - orthogonal to the Web's organization. However, by looking at the architecture and technologies of the Web you can see the areas where this multi-site capability can exist :
    • clients such as rich internet applications, browsers and browser extensions like toolbars
    • intermediaries such as the proxies that ISPs like Comcast operate
    • compound resource structure of current Web documents. Since each resource can be retrieved from different domains, information can leak between domains.
    Look for control or partnerships in these areas in the future.

    May 14, 2007

    Command Links

    Uh oh. I see a mess coming up...
    From Jakob Nielsen's Alertbox post on
    Command Links
    "Windows Vista introduced a new GUI widget for commands: the command link. Once something is in the system that people use on a daily basis, it becomes a de facto standard. Because they'll encounter them frequently in Vista, users will come to know and expect command links."


    From what I gather, Vista 'command links' appear to be glorified buttons for native applications, not 'underlined text' links on Web pages.

    However, Jakob continues with the following regarding web page links:
    To reduce confusion, link text should explicitly state that it leads to an action and not just to a new page. It's not enough to communicate this info in the surrounding text; users often scan Web pages for the areas they can act on. Thus, you should assume that most users will only read the link text. In fact, users often read only the link text's first few words, so it's important to start with a word (typically a verb) that indicates the action that results if they click the link.


    It appears Jakob is also implicitly approving the use of a link on a web page as an 'action'. From my reading, I think this is seriously wrong. I believe web page 'action links' need to satisfy the following two requirements

    • visually distinct from normal 'safe' links
    • syntactically distinct from normal 'safe' links


    The second point is important because of the large number of automated agents that traverse the Web through hyperlinks. Adding unsafe links into the web will cause confusion among both people and software agents.

    May 09, 2007

    It’s my Vineyard

    I found this blog about a couple that have purchased a vineyard in France - what a dream!

    It’s my Vineyard
    "Sitting in glorious sunshine on the terrace of The Restaurant du Pont having a delicious lunch we are reflecting that it is almost two years ago that we fell in love and bought Maison des Bulliats and its vines in Regnie, Beaujolais, one of the most idylic spots on earth."


    This reminds me of the book "The Olive Season" which is about an impetuous British actress that purchases a run down villa and Olive 'garden' in the South of France. We found that book in the apartment we rented while we stayed in France two summers ago. Interesting book - although the scattered personality and writing of the author can get exasperating - and I'm looking forward to following this French vineyard blog.

    May 08, 2007

    XML is in the House

    I ran across this page - Legislative Documents in XML at the United States House of Representatives: "The purpose of this website is to provide information about the ongoing work of the U.S. House of Representatives in relation to the eXtensible Markup Language" - and thought the host name was very cool - xml.house.gov.
    This page has pointers to DTDs and schemas for the governmental processes involved with making and amending laws. So the next time we need to form a more perfect union, we've got this going for us.

    And that page led to this page of a summary of floor proceedings of the US House of Representatives. It reads like a blog, only with less detail and there is no RSS feed.

    May 03, 2007

    Apollo, Silverlight, blah blah blah

    Looks like Hugh W is one of the few questioning the value of Silverlight and RIA for the Web.

    Flash, applets, Silverlight, Javascript -- the more you use them, the suckier your web apps are at exploring the web information space. I don't think it has to be this way, but it takes a design discipline few seem to have. These programming models are from the 80s. They have web APIs, but they're not web oriented. Programs end up as little desktop applications, not web apps. I don't see Silverlight changing that. It is good to have super expressive widgets -- hear hear. But if you're not pushing a bunch of hypertext down to my browser, you're not helping me explore the space.


    I agree with his sentiment - you might say that RIA is to user interfaces as RPC is to messaging interfaces : more is not better. There probably will be a few years of smooth looking but hard to use (and harder to re-use) applications, while we wait for people to re-learn the basics of usability. I can only hope folks read Nielsen's UseIt column.

    May 02, 2007

    Mathematica 6

    It looks like Wolfram Research has release a huge new release of Mathematica. I have only dabbled in Mathematica but I could spend all day playing and learning with it. It's hard to believe Mathematica came out almost twenty years ago!

    Something they've added recently - and apparently improved on - is server based computing and visualization. Imagine what could be done by putting this together with something like Amazon's Elastic Compute Cloud (EC2).

    See the Wolfram Blog for more details.

    May 01, 2007

    REST - it's inevitable

    A couple weeks ago I was having lunch with a friend at Amazon - who coincidentally used to work at MS with the data access team and is a brilliant architect and engineer - and we talked about REST and how he was helping use REST concepts in refactoring some core back-end services. I made the observation that I was never stressed about how long it has taken for REST to achieve common industry understanding and acceptance - I just said "It's inevitable". I didn't realize just how short it would be for inevitable to show up.

    It looks like REST has taken Redmond by storm. First I read that Mark Baker did some consulting with Microsoft (I missed the chance to have dinner with him when he was in town due to email snafu - major bummer). Then I read Dare Obasanjo's post that says "REST is totally sweeping Microsoft."

    We passed the tipping point quite a while back, but it's still good to see pragmatic architectural sensibilities take root finally.

    April 25, 2007

    Sparkly things factory

    Oh, this quote from Performancing.com is beautiful!

    "Snap's preview anywhere gizmo is ruining the reading experience for millions of people. Its intrusive, obstructive and unuseful in almost every respect and use case. The fact that so many big blogs are using it, big well respected blogs, does not mean that it's useful, it just means that they, like most bloggers, have all the self restraint of a magpie in a sparkly things factory."

    April 24, 2007

    Amazon.com Widgets and hypertext

    It looks like Typepad and Amazon are collaborating to help bloggers link to Amazon products. They have three widgets outlined, and one of them - the quick linker widget - uses custom attributes on an HTML anchor tag to make it easier to reference a set of products. The 'old school' would just use a URI, but those are hard to construct, hard to type and the wizards slow people down, blah blah blah. This not-quite-micro-formats approach is more understandable and more forgiving for hand-crafted markup. They define a new 'type' attribute for the anchor with the value "amzn". Then there are several more attributes like 'search' or 'category' or you can create a direct link via an 'asin' attribute. I assume some snippet of javascript would scan the page after it was loaded and construct the URI on the fly and set the href attribute on these anchors.
    This is a creative solution to the "how do you construct a URI" problem, but it does leave spiders out in the cold, breaking the hyperlinking which defines the Web.

    However, there is a simple approach they could use which would make a fully declarative and locally described document work and continue to allow auto-discover of hyperlinks to work - add a 'meta' tag to the head of the document with the URI template that corresponds to the 'type' attribute. I think the URI template proposal may need to do a bit of work related to optional or conditional patterns, or the URI template could stay static and the type="amzn" could change to be type="amzn-direct-link" or some other more qualified value.