Wednesday, June 24, 2009

Leave the Gun, Take the Cannoli

I grew up in New York City and spent enough time in Jersey back East to have witnessed a couple interesting brawls in my life (I even had neighbors who dug holes for a living or worked in waste management) but it’s been a while since I’ve seen anything like the recent scuffle among industry analysts and vendors regarding the recently published ParAccel TPC-H benchmark. Maron!

It all started innocently enough two days ago when Merv Adrian, BI industry analyst emeritus, published the news in his blog titled “ParAccel Rocks the TPC-H – Will See Added Momentum”.

Now, it’s not every day that a vendor publishes audited TPC-H benchmarks (“audited” being the key word, as the process runs around $100K from what I understand). Very few companies besides the Big Three have the deep pockets and technology prowess to accomplish that. Furthermore, ParAccel did its benchmark based on 30TB which isn’t exactly a small chunk of data. And so Merv made the point that, at the very least, the news should certainly help put ParAccel on the map. To quote him: “This is a coup for ParAccel, whose timing turns out to be impeccable”.

Immediately, this was picked up by no other than Curt Monash, BI analyst to the stars (and I say that quite seriously), who happens to despise the very concept of TPC benchmarks for reasons he clearly outlines in a recent post entitled “The TPC-H benchmark is a blight upon the industry”. To pull a couple of money-quotes from the site:

“...the TPC-H is irrelevant to judging an analytic DBMS’ real world performance.

“In my opinion, this independent yardstick [the TPC-H] is too warped to be worth the trouble of measuring with.

“I was suggesting that buyers don’t pay the TPC-H much heed. (CAM)”

“TPC-Hs waste hours of my time every year. I generally am scathing whenever they come up

Now, notwithstanding the TPC-H issues, I think Curt will concede that he doesn’t particularly appreciate or trust ParAccel as a company either as the following statements will show:

“I would not advise anybody to consider ParAccel’s product, for any use, except after a proof-of-concept in which ParAccel was not given the time and opportunity to perform extensive off-site tuning. I tend to feel that way about all analytic DBMS, but it’s a particular concern in the case of ParAccel.

“I’d categorically advise against including ParAccel on a short list unless the company confirms it is willing to do a POC at the prospect’s location.

“The system built and run in that benchmark — as in almost all TPC-Hs — is ludicrous. Hence it should be of interest only to ludicrously spendthrift organizations.

“Based on past experience, I’d be very skeptical of ParAccel’s competitive claims, even more than I would be of most other vendors’.

The combination of published TPC benchmarks and the originator of the benchmark seem to have created what Curt himself refers to as “the perfect storm”. To say he doesn’t like either would be a gross understatement :)

Both blogs immediately started getting “opinionated” comments from the public at large, including ParAccel’s VP of Marketing Kim Stanick and a gentleman named Richard Gostanian who may or may not be connected to Sun Microsystems (depending on which Twits you read). Sun supplied the hardware for the ParAccel benchmark. To cite a couple quotes from the comments, Richard Gostanian responds:

“Perusing your website, I detect a certain hostility towards ParAccel.” – (No kidding!)

“Indeed you do more to harm your own credibility than raise doubts about ParAccel.

“…TPC-H is the only industry standard, objective, benchmark that attempts to measure the performance, and price-performance, of combined hardware and software solutions for data warehousing.

“So Curt, pray tell, if ParAccel’s 30 TB result wasn’t “much of an accomplishment”, how is it that no other vendor has published anything even remotely close?

Then Kim Stanick says: “It [TPC-H] is the most credible general benchmark to-date.”

And an anonymous reader chimes in:

“After reading Curt’s post about ParAccel and Kim this is obviously personal…I wonder why the little fella didn’t have a fit over Oracles 1TB TPC-H? Check his bio. He consults for Oracle.”

To which Curt replies (among other things):

“As for your question as to why other vendors don’t do TPC-Hs — perhaps they’re too busy doing POCs for real customers and prospects to bother.

Ouch! This nasty sudden melee took me by surprise at a time when I was considering blogging about the whole TPC-H system for analytical engines anyway. I’ve wondered for quite a while whether or not publishing such metrics actually helped “new breed” startups like ourselves from a marketing and sales standpoint. Given the high cost and resource drain, what’s the return on this investment? What’s more, I have yet to meet a prospect or user who either cares of knows about TPC-H benchmarks. So far, the only people I’ve ever seen show any interest in the matter are venture capitalists and investors which tells me right there that something is amiss (or, maybe that’s why the small players take the plunge, I don’t know).

As some of you may know, XSPRADA is a recent member of TPC.org alongside other industry startups like Kickfire, Vertica, ParAccel and Greenplum. Numerous other startups in the same category are not members. They don’t seem to fare any worse. Furthermore, as best I can tell, even some existing members (namely Greenplum or Vertica) don’t publish audited benchmarks. Yet clearly these two vendors don’t seem negatively affected by the lack thereof.

Although we at XSPRADA have conducted TPC-H benchmarks (and continue to do so) internally, we have never attempted to get them audited and published. If a prospect asked me about it, I would recommend we help him run those benchmarks in-house on his own hardware anyway! Even if we had $100K to blow on getting audited benchmarks, I’m not sure it would make sense to pursue.

I’m usually a pretty opinionated black & white guy, but with respect to this TPC-H business, I tend to centerline. Strangely enough, I identify with both sides of the argument. One the one hand, I don’t believe the benchmarks to be totally useless. Having been involved in generating our internal results, I can vouch for the fact that it takes a lot of tedious work and kick-ass engineering to even complete the list. By any stretch of the imagination, this is not a small inconsequential feat. Doing so on anything above 10TB is, in my opinion, nothing to sneeze at. If nothing else, being able to handle the SQL for all 22 queries is a decent achievement. And then of course, there’s the notion that even “trying” to do it is noble in itself. In that sense I tip my hat to the small guys who pulled it off.

On the other hand, I don’t feel the benchmarks are holistically useful for evaluation purposes. As a prospect looking at several vendors, they might figure in my check-list but not more significantly than others I consider more important. Namely: how easy is the product to work with, what resources does it consume (human and metal), how does it play in the BI ecosystem as a whole (connectivity), and last but not least, what kind of support and viability will the vendor provide? I’m a little weird that way. I tend to evaluate companies based on their people over most everything else. But that’s just me.

At the end of the day (and everyone does seem to agree on that), what matters are onsite POCs. Nothing can beat running your own data on your own metal. I want a vendor to hand me the keys and go “ok, have a good ride, call me if you need anything” and mean it. BMW sells cars this way. Enough said. This is what I drive to when helping people evaluate our offering.

It remains to be seen how much of this brouhaha will benefit ParAccel in the long run. They say there’s no such thing as bad publicity. If they end up getting recognition and sales from it, then they have chosen wisely, and no one can take that away from them. Personally, I wish them the best. I believe the more numerous we are in this upstart game, the better it is for us, and more importantly, for our customers. So I say leave the guns, and take the cannoli.

Tuesday, June 23, 2009

In-House or SaaS? How About Both?

Chuck Hollis just penned another interested blog post about the economics behind private clouds for the enterprise. There is a lot of talk about on-premise versus on-demand SaaS these days in the BI community (and when I say SaaS I mean either private or public).

From a financial standpoint, the two models are fairly well established. Basically, on-premise is budgeted as capital expenditure, while on-demand is budgeted as operational expenditure. Much like the difference between buying your TV and paying for your electric bill on a monthly basis, or the difference between buying and leasing a car.

With on-premise you buy a lot of expensive stuff and it depreciates (and loses value) over time. With on-demand you rent a service for a short critical amount of time as needed. In many ways, the parallels are strikingly close to hiring in-house software developers versus outsourcing to consultants. The business case for either direction is easily conceived.

In-house developers are a long-term investment. They will learn the business and be allocated as needed on a per-project basis. The project is likely to be long-term. Much like capital equipment, they also need to be “upgraded” periodically – namely allowed and encouraged to keep up with technology so they remain productive and far-sighted. They also need to be provisioned with tools and resources to do their jobs effectively. The most progressive shops understand that. Not unlike expensive heavy metal and software licenses, they also tend to get worn out or outdated with time. And they’re expensive to replace and renew.

Consultants (either remote or in-house) are a quick-fix solution, usually applied to a pressing problem when in-house resources are either not available or not competent to handle the pressing business need. They’re expensive, but they (hopefully) get the job done quickly, get you the answers you need when you need them, and ride out into the sunset. Like on-demand offerings, they can be shutdown at will, but they also carry vendor-lock risk.

Having been in the software business for two decades, I’ve been on both sides of this fence numerous times. In my experience, the best shops implement a hybrid approach with strong internal cores supplemented as needed by top-notch “gun slingers”. In the best of cases, I’ve seen synergy and knowledge transfer occur between the two entities (when the politics were right) with significant benefit to the enterprise.

I think the same thing will probably occur in the BI space. I would imagine large shops will probably have both in-house staff and equipment, backed up by quickly ramped up SaaS offerings dedicated to what I call “transient data mart needs”. I could be wrong about this, but hybrid approaches (in business and technology) are usually more flexible and not necessarily conflicting. They can also feed off each other in positive ways.

I don’t believe the choices are going to be 100% on-premise or 100% on-demand. The trick for the CIOs out there is going to be determining which projects and which needs are better served by internal (strategic) or external (tactical) solutions in an agile way. In that sense there is really no “battle” between the two approaches. They should be considered complementary parts of an intelligent BI strategy toolbox.

This ACID Leaves No Bitter Taste

In my previous post below I received a comment (question) from Swany about inserts in the XSPRADA database engine RDM/x. Specifically, he (or she) asked:

“What happens if there is an error during the SELECT .. INTO? Are such inserts ACID or will I get partial data in a table if the system crashes?”

This is of course an excellent question, and I thought it was worth addressing on a wider level beyond just incremental loads. To place this in context, recall that ACID properties are defined as a set of rules pertaining to transactional database management systems defined as Atomicity, Consistency, Isolation and Durability. If a transactional database does not meet these conditions, it is not considered “reliable”. I won’t bore the reader with yet another ACID definition. Suffice to hit Wikipedia for a reasonable description.

Now the first question of course is whether analytical database engines supporting OLAP style work can or should meet the same criteria as classical transactional OLTP systems. By definition, analytical systems are biased for read access and updates are supposed to be rare, but inserts certainly occur as incremental loads are performed on warehouses (and data marts) at various intervals (from hours to weeks typically). In either case, an analytical engine clearly needs to handle data changes in an ACID way or data loss and corruption can occur. Similarly, data value and integrity need to be protected (locked) from concurrent (possibly conflicting) access patterns. Internal database structures are vulnerable to corruption during these transactions. So how does XSPRADA technology handle these issues?

The answer, not surprisingly, lies in XSPRADA’s “magic sauce”, namely, the mathematics of Extended Set Processing (XSP). To appreciate this, one needs to understand that all entities inside the XSPRADA mathematical model (ie: tables, rows, fields etc.) are represented as extended sets. And all extended sets by definition are immutable. This means updates to the system are implemented by creating additional extended sets, and the original ones are never mutated or deleted by subsequent processing. This ensures that data sets in the system are never corrupted or worse, deleted by mistake.

Internally, the XSPRADA data model is fully contained in a “universe” of extended sets. This is the set of all sets. Sets in this universe are related to each other via algebraic relations (hence the “relational” part of Relational Data Miner or RDM/x). Depending on the state of the system at a specific point in time, sets are either “realized” (materialized) to disk, or “virtual”, meaning they have an internal mathematical representation defined by algebraic expressions involving other sets, but no physical existence. (This has repercussions concerning “materialized views” which I’ll attempt to discuss in a future post).

So when sets are modified by internal operations, both new and old sets remain in existence within the universe. This means updates and inserts never actually change information, but only add to it. To complete the transaction, RDM updates the universe metadata to include knowledge of the newly created sets (if any) along with the algebraic relations linking them to their original brethren. Genesis is maintained. This is crucial because, unlike in conventional DBMS systems, the original data never needs to be re-generated (or created) to achieve rollback. The universe is only updated once all operations have completed successfully. If an error occurs, no harm no foul, as the previous state of the universe was maintained and still exists. In essence, the INSERT, UPDATE, DELETE functionality of the XSPRADA database is merely a logical emulation of conventional DBMS DML. Each of these statements internally results in an additive activity. In fact, UPDATE is internally implemented as DELETE+INSERT. So to recover from a failed set of operations (a transaction gone south) the system simply deletes any incomplete sets and does not update the universe metadata! This mechanism enforces atomicity and consistency natively without any need for additional programming or complexity.

On the isolation front, lists of in-process and pending operations are maintained in dynamic pipelines for each extended set. The system examines these pipelines and algebraically identifies any potential conflicts between operands. Again, the mathematics allows this to occur natively. So if the results of pending operations do not affect in-process operations, the system executes them concurrently and immediately. Conversely, if the mathematics identify a potential conflict or deadlock, pending operations are queued until conflicting running operations have completed.

Durability is the last remaining condition. The system maintains all realized sets in persistent storage (disk drives). Although sets or parts thereof can be (and often are) kept in cache, any complete set also has an image on disk. This prevents system failures from affecting the durability of realized sets. When the system restarts, it also restarts any operations that were executing at the time of failure.

For all these reasons, XSPRADA technology is actually superior to conventional database mechanisms for enforcing ACID, as the enforcement is inherently “built-in” via the mathematics underlying the system at all times. As I mentioned in the last post, the engine is also time-invariant, meaning it can always be queried at a given time point in the past. The ability to do this without any external programming or internal modeling is significant. One use case that immediately comes to mind (to me anyway) in wake of the recent Wall Street disasters is being able to ask a financial database to yield answers as if it were being queried months or years ago. Imagine being able to roll back time to analyze or audit results that supported past decisions and the people who signed off on them. What a concept!

If you'd like to take the XSPRADA database out for a spin (and pull the plug in the middle of using it just to see what happens ) feel free to get the trial bits from our website.

Monday, June 22, 2009

Can Rover Roll Over Too?

In this post, I want to stay in gear-head mode and "demo" a couple of neat tricks from the XSPRADA RDM/x analytical engine puppy. I’m going to address quick prototyping capabilities (thanks to schema agnosticism), incremental inserts, and time invariance.

We’re going to do this with some really simple tables and data structures just to highlight the concepts behind the engine’s functionality. So suppose we start with a CSV table called ‘demo.csv’ consisting of the following rows (NOTE: the file must be terminated by CR-LF):

1000,12,"now is the time"

1001,45,"for all good men"

1002,76,"to come to the aid"

Now we want to load this into RDM/x so we use the XSPRADA SQL extension CREATE TABLE FROM as follows:

create table demo(id int,c1 int,c2 char(128)) from "C:\ demo.csv"

Then we do a SELECT * on the table just to make sure all is well (for example, using QTODBC or any other ODBC-compliant SQL front-end tool) and we see:

1000 12 now is the time

1001 45 for all good men

1002 76 to come to the aid

And if I look at my schema in the QTOBDC object browser I see this as expected:

So this is fine and dandy when you know the actual schema of your data but now what if you don’t or what if you’re not sure or what if you're out to determine what the best schema might be given your application?

Ok well what we can do it simply load everything as text types. So now we do:

create table demo(id char(128),c1 char(128),c2 char(128)) from "C:\ demo.csv"

The original ‘demo’ table is overwritten with the new schema which now becomes:

And now we can play around with that until we’re satisfied. Assume we suspect the optimal way to work with this data might be using INT for the index, a double precision type for the c1 column, and an 80-char VARCHAR for the last field. We might take our existing table and “flip” it into a new schema (called flipped) as such:

select cast(id as int), cast(c1 as double precision), cast(c2 as varchar(80)) from demo into flipped

Note how we convert the schema into a new one directly into a new table on the fly. Now if we select from flipped we see:

1000 12.0 now is the time

1001 45.0 for all good men

1002 76.0 to come to the aid

And the schema for the new table, as expected, is:

Now if we needed to do some querying on numerics for C1, we could. Note we didn’t even have to explicitly create or schema the ‘flipped’ table. RDM/x took care of that automatically, making these types of operations ideal for quick prototyping. Incidentally, we could have done this directly on the existing table as well. This type of flexibility is pretty cool and allows you to “play” with the data in a trial and error mode.

Now suppose we come across some additional data and we wish to INSERT this information into our existing database. This is typically what happens with daily incremental into a data warehouse. On a periodic basis, chunks of data are added to (typically) fact tables. Our incremental CSV looks as such:

2000,132,"to be "

2001,465,"or not to be"

2002,786,"that is the question"

There are several ways to handle this using RDM/x. The most direct one is to simply use the XSPRADA “INSERT INTO FROM” SQL extension as follows:

INSERT INTO demo FROM “c:\demo_062209.csv”

In this case RDM/x is told “hey, there is more data being presented to you so, algebraically, union it with the existing data”. As RDM/x does not “load” data in the conventional sense of the term or support bulk loading (as it doesn’t need to), additional data can be presented virtually in real-time with little or no effect on the ability to query the system simultaneously.

This is what’s called “non-disruptive live updates” or concurrent load and query capability. Initially, users are a little surprised that bulk inserts are not supported. But in fact, all that’s needed is “dropping” the incremental data somewhere on disk and telling RDM/x of its existence. RDM/x can ingest this information as fast as it can be written to disk.

Another way to do this is by loading the incremental “chunk” into its own table as such:

create table demo_062209(id int,c1 double precision,c2 varchar(80)) from "C:\ demo_062209.csv"

And then doing

select * from demo union select * from demo_062209 into demo

A bit of subtlety there: when you do this, you’re effectively “updating” the existing demo table with the unioned results of the incremental. So you take the old demo table, add the incremental, then flip that back into the original demo table. You could have saved the “old” demo table first before doing this as such:

select * from demo into demo_current

Notice how these “dynamic” tables really behave like variables in a loosely-typed programming language. They are conceptually related to conventional database “views” (although RDM/x does support view objects in the database but from a purely semantic way). As a matter of fact, all tables and views in RDM/x are essentially the same thing and materialized (to disk and/or memory) on a JIT basis anyway.

Of interest here is that you can always recover past instances of any data RDM/x via a nifty little feature called “time invariance”. This feature is unique in the world of databases to the best of my knowledge. Essentially, you can query RDM/x much like a “time machine”, asking it to yield results as if the question were being asked in the past.

So suppose you had inserted your incremental into the ‘demo’ table by mistake and suppose your insert had occurred at a given time such as '2009-06-22 15:13:52.775000000' (and you can tell this if you are logging your SQL statements to the database using the usrtrace option of the RDM/x ODBC driver). You can always recover the state of ‘demo’ at the time by issuing:

SELECT * FROM demo AT TIMESTAMP '2009-06-22 15:13:52.775000000' into recover

Now your ‘recover’ table will contain ‘demo’ exactly the way it was at that time.

The implications are far-reaching because you can essentially always query the RDM/x database at any point in the past. RDM/x never deletes information UNLESS the information has been explicitly deleted using a DROP TABLE or DELETE type of DML statement AND the garbage collector kicks in. Short of that, the database maintains information and integrity through time.

These are just a couple of nifty tricks from one original database engine called RDM/x.

Friday, June 19, 2009

Choose Wisely Grasshopper

If you’re been tasked with implementing or supplementing business intelligence offerings in your organization lately (read: a data mart, for example), the choices have gotten quite a bit more complicated. A while ago, you would have your pick of three or four vendors. If one of those vendors was already in-house handling your transactional systems, guess what, he was likely to also handle your warehousing needs. Or at least try to.

Nowadays, the decision path is a little more involved because the BI world is no longer ruled by a small oligarchy (namely Oracle, Microsoft and IBM). The proliferation of “new-breed” analytical engine vendors has greatly expanded a buyer’s options. There are around twenty-some players in this field now. Not only that, but delivery options have expanded as well.

Nowadays, you can get BI delivered and running in-house sitting on commodity or custom hardware. You can buy canned appliances. You can go proprietary bits. You can go Open Source. Or you can tap the “cloud” with an on-demand subscription-based model. You can pick a columnar vendor, or a row-based one. You can choose an MPP architecture, or an SMP implementation. You can even stick with the “gorillas” if that makes you feel better (and money is no object). These different paths are all strategically different both technically and economically. The modern BI strategist is compelled to choose wisely in an unforgiving economy where failure is no longer an option (this time they mean it).

So for the sake of argument, I’m going to assume the following buyer profile:

  1. Money is _definitely_ an object.
  2. IT resources are non-existent, limited, not available, or not inclined to help.
  3. A lengthy proof-of-concept (POC) cycle is not an option.
  4. Project timeline is measured in weeks not months or years.
  5. C-level people want to see incremental results starting today.
  6. A single DBA is left standing in your organization (but next week, maybe zero).
  7. Your ass is on the line.

I think this describes a fairly common scenario these days. Faced with such odds, I think most people opt for the path of least economic and implementation resistance. Most people don’t have $2-6M hanging around including staff and a six-twelve month window to implement a full-blown Oracle, Microsoft or IBM solution. These days are simply gone (good riddance on that). It would appear at first glance that the only remaining alternatives would be open source (OSS) or on-demand SaaS software.

Now, OSS is attractive from a cost basis, as most freebies tend to be. Yes, you do pay for support if pulling down “enterprise” versions of the software but in many cases, some buyers get away using the free versions for a while, at least for quick POCs. In some cases, the free versions are either limited or incomplete in functionality (like InfoBright, for example) so that could be a “gotcha” depending on your application needs. Similarly, non-enterprise versions usually depend on community for support. If you’re in a jam and need serious dedicated support on a moment’s notice, you’ll need to pony up for an enterprise version or wait until “the community” comes up with an adequate answer to your problem (if ever).

Additionally, OSS does have hidden costs, not the least of which are installation, setup, configuration, and maintenance. But, if you happen to be a Linux shop and have enough LAMP developers on staff with sufficient expertise and time, and your management happens to be accepting of the whole OSS concept, it might just be a viable option.

Another quick way to get up and running quickly is to go the Cloud (SaaS) service route. In that scenario, you pay a monthly fee to access a BI platform in the cloud. This hands-off approach is certainly attractive in many cases provided data volumes and security restrictions do not get in the way. Shlepping 10-100TB of data offsite is not something most people consider yet. But, for smaller data sizes, SaaS is certainly an option. Vertica, Kognitio and Aster Data come to mind as the latest new-breeders to provide cloud-based services to customers (either on proprietary or public cloud platforms like EC2). There is a flurry of other on-demand BI players as described in several of my past posts. Of course, the downsides there tend to be upload time, limited functionality and vendor lock-in.

Now if I can get on my soapbox for a minute, I’m going to pitch a third option. At XSPRADA, we’ve developed a 5MB high-performance analytical database running as a Windows service on commodity hardware and Server 2003 or 2008 x64 operating systems. You can install this puppy internally (your data stays nice and safe in-house) in about 2.5 minutes including ODBC drivers. You then point it at your CSV data on disk and start firing off queries immediately. The more you ask, the faster it gets. And you can show results in minutes, not weeks or months. No cubing, no indexing, no pre-structuring, no partitioning, none of that nonsense. That’s the bottom line in a nutshell. I promise you business intelligence is not the rocket science so many folks make it out to be! For those who might be hesitating between open source, SaaS or much more expensive in-house solutions, I think this is a pretty unique proposition. There really is nothing like it anywhere else on the market, and it just so happens we have a 30 day trial going on right now if you visit our website J

Thursday, June 4, 2009

Yabadabadoo and off to Austin!

I’m headed out to Austin again tomorrow. Last I checked it was 94F over there with 60% humidity. Guess I won’t be going on my daily power walk for a little while lest I keel over. That’s OK though, I’ll get my workouts at Bone Daddy’s. I’m jazzed up about the trip as usual. I love that town. It has a relaxed, genuine, homey feel I have not found in many other US cities. When people ask “how ya doing?” they actually mean it and expect an answer!

I have prospects, clients and partners to visit. Now that our pre-release RDM/x software is available, I’m getting to see how people actually use these bits in real-life situations and that’s fairly exciting for many reasons.

Product claims are one thing. But it’s quite another to actually see people’s reaction when they discover the software. When folks install our software the usual reaction is “huh? Is this all there is to it? What did I miss?” – The answer is “you missed nothing” – RDM/x deploys and installs as a 5 Meg Windows Service. Read my lips: a 5 megabyte Windows Service EXE is all you need to query TB-level data sets faster and easier than you’re used to. Install the bits, present data, connect to ODBC driver, and ask questions. Done. Naturally, given users’ past experience and expectations with classic analytical engines, it’s hard to swallow initially. And speaking of shockers, here a doozy of a benchmark you can find on the TPC website.

That’s right folks, $6 Million dollars with Godzilla metal and full-blown enterprise software will buy you queries into a mere 1TB of data. Or, you could install and start a 5MB Windows Service on a $20,000 server.

What else am I picking up? People are incredulous about the fact that there’s no need for partitioning, indexing, pre-structuring, worrying about rigid schemas and data models. In one case, I was actually able to load 18GB of data by pulling everything in as strings. These particular data sets were CSV clickstreams for a web analytics application (conveniently, they’re often text log files which suits us well as CSV is the only format we currently ingest). The ETL attempts had become tedious due to poor data quality and uncertain schemas. It wasn’t clear initially what the best schema might be for the type of data at hand so experimentation was needed. So we took more of an ELT approach and just DDL’d everything in as text. We then quickly flipped internally to another schema, casting columns as needed on the fly. And if that schema turned out to be less efficient or interesting than originally thought, we just flipped the whole dataset to another one in one line of SQL. This type of just-in-time flexibility in presenting and transforming data is seldom (if ever) found in more ponderous solutions. For web analytics type of applications (where DQ is typically poor and schemas fairly dynamic) this is a huge competitive advantage.

Another sweet spot is OLAP cubes. Or lack thereof I should say. In typical OLAP applications, building cubes is lengthy, complicated and expensive (as anyone who ever footed an analytical consultant’s bill will attest to). Re-aggregating or adding dimensions is also painfully slow and tedious as data volumes increase (and they always do). A lot of times these processes are setup to run nightly as batch processes. If you’re lucky and if you’ve done your engineering properly (lots of ifs), you show up in the morning and it’s done. But you can’t add a dimension in real-time without impact on the entire system. And you need to plan for your questions up-front, as in “what shall I slice & dice this year?” With RDM/x, you don’t need to setup dimensions explicitly. The very act of querying in a dimensional way alerts the engine to that effect and it starts building “cubes” internally and aggregating as needed, optimizing for specified dimensions. Then adding a dimension is painless – all you need to do is actually query against it more than once. With RDM/x, once is an event, twice is a pattern.

This “anticipatory” self-adjusting behavior is actually at the heart of RDM/x. This is one area where the product distinguishes itself from the competition. Because the software continuously re-structures both data and queries as more and more questions come in. This means you see response time decrease as quantity and complexity of queries increase. This is fundamentally different from conventional database behavior where query response time “flat-lines” up to a given number (not including load time) and pretty much hovers there until and unless an “expert” can optimize either code or configuration (if possible at all). And when you add to or complicate your query patterns, performance suffers.

The RDM/x response-time profile is more of what I call the “dinosaur tail” effect: big initial hump, then exponential taper down over time. Here’s my Fred Flintstone rendition of the effect using MS Paint (nostalgia?):

It is always a blast observing people as they first witness this unique behavior. It’s like discovering a new life form J -- And I’m looking forward to repeating the experience in Austin next week and throughout the country soon enough. If you care for a head-start, pull down the bits from our website and give it a whirl!

Monday, May 25, 2009

Data Warehousing for Dummies

Every so often I’ll be talking to someone implementing our solution at a customer or POC site and the question comes up “what’s so special about your database anyway?” and “So, is it really different from MySQL, SQL Server or Oracle?” or “I don’t understand why your database talks SQL since it’s not a normal database like Oracle”. Better yet: “What else do I need to do after installation to get this working?” Usually these questions come from non-LOB folks who are tasked with implementing a particular solution using our product. Typically these people tend to be experienced software developers or DBAs. People who, as Joel Spolsky likes to say, are “smart and get shit done”. The type of folks you can throw a problem at and say “Ok, go solve it using this new tool.”

But as often happens in large organizations, they may not have been briefed fully by management on the features/functionality of the new tool needing evaluation. Or maybe this is their first exposure to BI. They may also never have encountered or worked with an analytical database product. I know that, several years ago, if you’d asked me what the difference was between OLTP and OLAP I would have blurted something like “one if for transactional stuff, the other for reporting” and been in the right ballpark but no cigar.

So when these questions come up, I am always ecstatic to be able to share what I’ve learned with the guys in the trenches doing the real work! The first thing I do is give I very general view of the differences between transactional (operational) and analytical use cases. Then I’ll try and give a 30,000 foot picture of data warehousing and its history. I‘ll mention Kimball and Inmon, of course, then several books and a series of blogs, websites and youtube videos for further exploration.

But this weekend, I discovered the holy grail of data warehousing 101. I was hanging out at my local Borders looking to trade my 40% off coupon in exchange for yet another good data warehousing/BI book when I noticed the yellow “Data Warehousing for Dummies” on the bottom shelf. Obviously I couldn’t resist picking it up, especially since I have yet to meet anyone remotely “dumb” in this business.

To my great surprise, I noticed the author was none other than Tom Hammergren, the owner of Balanced Insight, one of the top BI Software Innovator firms in the country. To say that Tom is a warehousing and BI guru is an understatement. This much I knew. But I had NO idea he was also an accomplished writer who could present this complicated subject in clear, simple terms anyone can understand and relate to! From now on, whenever someone asks me for the quick low-down about BI and data warehousing, I’ll be referring him or her to Tom’s book.

Now to answer the above questions about our own product RDM/x. There is nothing magical about our analytical database, at least from a usage standpoint. RDM/x talks and walks like any other database product out there on the market using ODBC. The magic is on the inside, certainly not in the interface (thankfully) as it supports a significant subset of the SQL-92 standard (minus TCL and DCL). Is RDM/x “really” different from SQL Server, Oracle, MySQL or DB2 though? You bet.

RDM/x is designed for data analysis, not transactional processing. As a matter of fact, RDM/x is the smallest, nimblest, on-premise solution available on the market that will let you query terabytes of data in minutes from installation. And that’s why I believe most people are a little confused from the get-go. Because they’re used to large footprint multi-module database clients with 500 page installation and setup manuals, followed by complicated tuning and optimization techniques involving indexing, partitioning, and all that “good” stuff. When they see a 5 megabyte piece of software installing as a Windows Service, ready willing and able to handle queries on giga or terabytes of data within minutes, they think they’re missing something. How can BI be this simple? Well it can. The proof is in the pudding and since we’re allowing you to download a fully functional 30-day evaluation from our website effective now, the best I can do is recommend you take me up on that assertion by visiting our website.

Thursday, May 21, 2009

Running RDM/x on Amazon EC2 is DICEE!

I want to keep this post fairly brief because there is so much stuff going on at XSPRADA lately that I find myself pressed for time from 6AM to midnight on a typical day which usually also includes weekends, but that’s the price you pay for building a revolution. Ask Fidel, he knows.

First of all, I finally had time last Sunday to record a screencast explaining how to install and setup our RDM/x software.  In the process, I discovered that Camstudio and Microsoft Windows x64 decoder were my friends, reducing a 1.2GB video to 35Megs (phew!).

Second, I want to talk about my recent epiphany with EC2.  Early this week I decided to see if I could install and run our Windows-based database engine RDM/x on some sort of cloud platform because I don’t think anyone in their right minds in enterprise software can afford to ignore this trend any longer.  My purpose was certainly not to setup a full-fledged production system up there, but rather to setup a quick and dirty demonstration system so people could either duplicate or use it on the fly to test-drive our software, for example.

After poking around a bit I settled on Amazon’s EC2.  They seemed like the only “big-time” player supporting WinTel boxes (our software runs on 64-bit Windows Server 2003 and 2008) and they have enough credibility and market “karma” at this point to alleviate most basic concerns about reliability and security.  So after checking out possible configuration tools and hitting our CEO up for some plastic, I signed up for EC2 and started exploring this brave new world.

It turns out EC2 is really several “platform” components comprising: the actual EC2 O/S instance (a VM blade) known as an AMI (Amazon Machine Instance), persistent storage called EBS (Elastic Block Store) which presents as “volumes” you mount onto the AMI, and persistent (hot/cold) storage (also used for EBS snapshots) called S3.  There’s also a queuing system called SQS but that didn’t enter my mix.   Confused yet? It’s not that bad once you get used to it J

For the configuration tooling, you can use command line tools (which I suspect most *NIX/LAMP people prefer), a FireFox pluggin, or the web-based AWS (Amazon Web Services) Console.  I used both of the latter to compare.

I brought up one of the standard Windows AMI as a Windows 2003 R2 64-bit datacenter server.  I actually tried two different instance types. One extra-large standard with 15GB of RAM and 4 cores, and one large high-CPU with 7GB of RAM and 8 cores.  I found better performance on the 4 core box with twice the RAM so I ended up sticking to that one.

For credentials, you need to generate a key-pair, and then plug in the private part into a dialog box which then spits out an admin password for the new instance. You then connect remotely to your instance using Remote Desktop Connection (or SSH if you’re talking to a *nix instance).

Right off the bat my instance came with four attached 500GB “hard drives”.  These volumes are more like flash drives I think. This is not persistent storage but it’s pretty darn fast.  For “real” storage you need to create Elastic Block Storage (EBS) “volumes” and attach them to your instance.  So I did just that and slapped 4 additional 500GB drives to my box, and then converted each drive to an NTFS mount point (because this is best practice for our particular application). Unfortunately, I extracted a maximum 22-25MB/sec I/O to and from these volumes.  I had read somewhere that these dynamic “block devices” were more like instant SANs but in fact, Amazon Silver Support (another pay-for service but well worth it if you ask me) stated the following to me in an email:

“Even though it presents a block interface, EBS isn't intended to be equivalent to fibre-channel SAN storage. The performance you should expect from EBS would more closely align with a NAS device. You can stripe several EBS devices together for higher I/O rates, but your rates will be limited by various shared components in the system, including the network between your instance and the storage servers. Larger instance types will typically see better performance than smaller instance types.”

“More closely align with a NAS device”.  Huh oh. That means gating at the NIC level.  The systems are clearly not setup for intense I/O data processing needs, at least not using the standard EC2 configuration models currently available.  Nevertheless, I was still able to do sufficient work with sufficient data to build a reasonable “functional demo” machine.  And that was my goal from the onset. Given this took me about 1.5 days to figure out, at an average cost of around $16/day (not including the additional Silver support fees) I am very impressed with this cloud platform to say the least and I’m sure Jeff Bezos is basking in the bliss of my endorsement :)

Quite honestly, this cloud business is no joke.  I haven’t seen, heard and felt such a buzz around a new “platform” in the industry since I got my hands on Windows 3.0 in the early nineties.  It was the same “oh my God” emotional feeling at the time, or DICEE as Guy Kawasaki likes to put it (Deep, Intelligent, Complete, Elegant and Emotive).

 

Saturday, May 16, 2009

Software that Sucks

Here’s a classic from the tech press that really caught my attention recently: 

“In the context of software, the word “Enterprise” has now officially come to mean software that sucks. Enterprise Software hit the nadir of suckitude (sic) at the launch of “Enjoy SAP”.  This is like the American Dental Association launching “Enjoy Root Canal”.  SAP is certainly an easy target, but let’s face it, “Enterprise Software” is generally a poorly integrated mess.  Working with Enterprise Software feels a bit like walking through an industrial landfill or an airport hangar.  Nothing is built to human scale.”

This was written on the SOA Center blog by no other than Software AG’s Chief Strategist Miko Matsumura.  His use of the techo-political term “suckitude” is one for the annals of our new post-TARP technology world.  If nothing else, the current situation seems to be facilitating proverbial “paradigm shifts” (namely, on-demand software) while encouraging more anti-status-quo “frank-speak” from industry figureheads.  I’m all for that. 

Because, notwithstanding all the pain, suffering and incertitude in the economy lately, one of the really brilliant consequences of this world-wide mess is that people are starting to say out loud what everyone’s been thinking silently for years.  Even in the sacrosanct enterprise software glass mansions, people who matter are starting to throw stones.  When major industry players start talking straight and using technical terms like “sucking”, you know the BS gloves are off.   I think established players, platforms and ways of doing business and thinking about customers are all up for questioning at this point.  Sunshine is the best disinfectant.

And speaking of gloves off, SAP and industry shifts, this old article from April 2008 refers to a slug match between two industry titans at the Churchill Club.  One is Marc Benioff from Salesforce.com and the other Dr. Hasso Plattner of SAP fame.  There’s a video of the exchange on Youtube.  I know it’s a long one, but I assure you it’s worth watching entirely if you care anything about the on-demand versus on-premise religious wars of late.

I am not going to comment at length on the video as anyone can draw their own conclusions, but I did want to point out what I consider some key points, and throw in a few gold nuggets.  

First, the body language between those two guys is simply priceless.  It is more than obvious from the get-go that they can’t stand each other.  You can catch the vibe even in that one picture in the article (and throughout the video).  Benioff’s looking away from Dr. Plattner constantly (he fidgets with his wedding band incessantly), and Dr. Plattner is reflective in his own world as in “why the hell am I here”.  To my amazement, at the end of the video, they both reveal that this is their very first in-person meeting!  Incidentally, one audience member does ask Dr. Plattner at the end why he accepted to do this.  His answer: “for the challenge”.  Not sure what that means.

Second, the verbal jousting between the two is fairly aggressive.  I don’t think these guys have much respect for each other notwithstanding their pseudo-polite claims to the contrary.  If you asked me whether Benioff hates Microsoft or SAP more, I’d be tempted to say SAP.

At one point Benioff states: "We have been passionate about moving obstacles out of the way of the old enterprise software companies.”  I guess this is one major tenet of the on-demand adepts.  Power to the users!  In my opnion, Dr. Plattner really  does buy the on-demand proposition but not “religiously”, and either way, he can’t say it in public.  He knows SAP screwed it up in the past.  I’m not sure he believes in SAP’s ability to execute such a shift internally.  And I bet he wouldn’t mind buying Salesforce outright with one check.  He implies as much several times but then claims he doesn’t want to get into a bidding war with Oracle.  Hogwash.

Throughout the video, both contestants score evenly, in my opinion, on the arrogance meter.  I guess they can both afford to be that way, but it does take a certain piece of the “human” side away from each.  For Dr. Plattner, I think the Germanic personality comes through more than genuine arrogance.  After all, he doesn’t need it at this point.  The guy built and ran a $40B company.  Enough said.  Benioff often has this “do the right thing” Google-ish “morality” in several other interviews and videos.  But when you watch him in action here, the only thing that comes out is ruthless self-convinced warrior (it’s no coincidence his favorite read is Sun Tzu’s The Art of War).  Although conviction and the ability to back it up is noble (and key to business success), I’ve always feared people immutably driven by their own dogma (mind you, I actually buy into on-demand big time).  But as my high-school math teacher used to say “you can never shelter yourself from a surprise”. 

Finally, as I was lauding “frank-speak” earlier, I did want to point out that Dr. Plattner uses the term “shit” several times during the exchange.  Initially, referring to Salesforce grabbing Dupont from them he states “Why did he win DuPont?  Because we had a shitty CRM system, and he had a much better one.”  Then later, referring to a customer still using code Dr. Plattner himself wrote: “…Shit! There is a customer in America still using the code I wrote.” Then referring to SAP’s earlier attempt at on-demand CRM: “…Shit, yeah!  It was better than our CRM on-demand.”   I find that endearing.

To conclude, if you truly want to understand the ongoing (and upcoming) battles between the SaaS and on-premise proponents of the enterprise software industry, you owe it to yourself to watch this video or, at the very least, pull down the transcript. And bring some popcorn!

 

 

Thursday, May 7, 2009

Tidbits and Check this Guy out

I read this article a couple of weeks ago and thought about one of our field test partners (telecom) because they had some political issues shipping us some data due to (very legitimate) privacy concerns – as in their CSO going “are you guys out of your f$##$ing minds?!?”.  

As it turns out, there are several data obfuscation tools out there on the market, including DMSuite’s offering as described in this article. I’m curious if most companies’ privacy policies make an exception for data that’s been altered by such a tool and if so, is there some sort of standard or certification these tools must meet?  If you know anything about that, I’d appreciate some insight.

I didn’t know until last night that there actually is a CIQP Certification.  What is CIQP you ask? Come on, get with the program!  Everyone knows what a Certified Information Quality Professional is!  There is a whole website dedicated to DQ as well.  I had never heard about this professional category.  If anyone reading this happens to be in that category and/or CIQP Certified, I’d love to chat with you and learn more about it.

For those of you who think the economy really sucks, your deduction is likely valid.  Nevertheless, BI and on-demand software market indexes seem pretty healthy to me as this article demonstrates.  My conclusion: I’d rather be in the “avant-garde” BI enterprise software sector than working for SAP or Oracle at this point J

I discovered Guy Kawasaki’s Entrepreneurial Lectures delivered at Stanford in 2003-2004 via this videocast series and sat there mesmerized listening to every single clip for hours.   

Guy (who now runs this blog and this company) successfully evangelized the Mac in the mid-80s and now runs a VC firm called Garage Technology Ventures.  His reputation and track record are legendary.  In the clips, he lectures young Stanford engineers-to-be on how to become successful entrepreneurs, change the world, and keep their soul in the process.  These are the points (or quotes) from his lectures that were etched on my mind:

  • Make meaning and make the world a better place.
  • Don’t write a mission statement, write a Mantra.
  • If your product is not unique and adds no value, you’re doing something stupid.
  • Don’t ask people [customers] to do things you wouldn’t do.
  • Be a Mensch.
  • Hire infected people.
  • Suck down.  The higher you go in the enterprise, the thinner the oxygen.
  • A milestone is something that increases the valuation of your company.
  • The valuation formula for a startup is: add $500,000 per engineer and subtract $250,000 per MBA.

All these points are perfectly in line with my personal experience.  But I’d never heard anyone formalize them in such an entertaining way before! And I don’t know that any comment can properly decorate any of these either. It’s one of those “you either get it or you don’t” kind of things.  It can’t be taught or inculcated by anything else than passion-driven experience.  

Tuesday, May 5, 2009

MAD About You.

I haven’t had a minute to sit down and blog lately. Our upcoming “pre-release” software slated for Cinco de Mayo has absorbed all my efforts. I just flew back from Austin, TX last week to participate in final engineering and release touches. At the same time, one of our major Defense prospects just re-activated a huge project so it’s all hands on deck at XSPRADA these days. I love it when a plan comes together.

Nevertheless, I recently picked up on a paper called “MAD Skills: New Analysis Practices for Big Data” referenced on Curt Monash’s blog by one of its authors Joe Hellerstein. Joe is a CS professor at UC Berkeley, if I understand correctly, and a consultant for Greenplum. I expected to read a lot of pro-Greenplum content in there but I don’t feel the major arguments presented are specifically tied to this vendor’s architecture or features per say. What makes this paper really valuable, in my opinion, is that it “resulted from a fairly quick, iterative discussion among data-centric people with varying job descriptions and training.” – In other words, it is user-driven (and not just a bunch of theoretical PhD musings) and as such, probably an accurate reflection of what I consider to be a significant shift in the world of BI lately.

Namely, that the “user” base is shifting from the IT type to an analyst and business person type. Note I didn’t say “end-user”, because the end user is still the business consumer. But what’s really changing is the desire (and ability) to cut out the middle man (read: the IT “gurus”), in essence, at every step of the way from data ingestion to results production. To put it in terms Marx would have liked, the means of production are shifting from dedicated technical resources to actual consumers J -- This is particularly true in the on-demand world, but probably pervasive throughout as well. I think the paper outlines and defines this change.

“The conceptual and computational centrality of the EDW makes it a mission-critical, expensive resource, used for serving data-intensive reports targeted at executive decision-makers. It is traditionally controlled by a dedicated IT staff that not only maintains the system, but jealously controls access to ensure that executives can rely on a high quality of service.”

The starting premise of the paper is that the industry is changing in its concept and approach to EDW. Whereas in the past an “Inmonish” view of the EDW was predicated on one central repository containing a single comprehensive “true” view of all enterprise data, the new model (and practical reality) is really more “Kimbalish” in the sense that the whole of EDW comprises the sum of its parts. And its parts are disparate, needing to be integrated in real time, without expectations of “perfect data” (DQ) in an instant-gratification world. This new model is premised on a MAD model: Magnetic, Agile and Deep.

Magnetic because the new architecture needs to “attract” a multitude of data sources naturally and dynamically. Agile because it needs to adapt flexibly to fast-shifting business requirements and technical challenges without losing a beat. And deep because it needs to support exploratory and analytical endeavors at all altitudes (detail/big picture) and for numerous user types simultaneously.

“Traditional Data Warehouse philosophy revolves around a disciplined approach to modeling information and rocesses in an enterprise. In the words of warehousing advocate Bill Inmon, it is an “architected environment" [12]. This view of warehousing is at odds with the magnetism and agility desired in many new analysis settings.”

“The EDW is expected to support very disparate users, from sales account managers to research scientists. These users' needs are very different, and a variety of reporting and statistical software tools are leveraged against the warehouse every day.”

Fulfilling this new reality requires a new type of analytical engine, in my opinion, because the old ones are premised on a model which, apparently, has not delivered successfully or sufficiently over time for BI. So the question beckons, what set of architectural features does a new analytical engine need to have in order to support the MAD model as described in this paper? And more importantly (from where I stand), is the XSPRADA analytical engine MAD enough?

To support magnetism, an engine should make it easy to ingest any type of data.

“Given the ubiquity of data in modern organizations, a data warehouse can keep pace today only by being “magnetic": attracting all the data sources that crop up within an organization regardless of data quality niceties.”

The “physical” data presented to the XSPRADA engine must consist of CSV text files. Currently CSV is the only possible way of presenting source data to the database. CSV is the least common data format denominator so this means the engine can handle pretty much any data source provided it can be morphed to CSV. Since the vast majority of enterprise data lends itself to CSV export, that covers a fairly wide array of data sources. How easy is it to “load” data into the engine? Very easy, especially since the engine doesn’t “load” data per say but works off the bits on disk directly. All it takes is a CSV file (one per table) and a SQL DDL statement such as “CREATE TABLE…FROM ” to “present” data to the engine. “The central philosophy in MAD data modeling is to get the organization's data into the warehouse as soon as possible.” – I’d say this fulfills that philosophy.

To support agility, an engine shouldn’t dictate means and methods of usage and must be flexible:

“Given growing numbers of data sources and increasingly sophisticated and mission-critical data analyses, a modern warehouse must instead allow analysts to easily ingest, digest, produce and adapt data at a rapid pace. This requires a database whose physical and logical contents can be in continuous rapid evolution… we take the view that it is much more important to provide agility to analysts than to aspire to an elusive ideal of full integration”

The external representation of data on disk and the internal “logical” modeling of the data changes dynamically based on incoming queries. The XSPRADA engine is “adaptive” in that sense and constantly looks at incoming queries and data on disk to determine the most optimal way of storing and rendering it internally. This feature is called Adaptive Data Restructuring (ADR). On the flexibility side, the XSPRADA engine is schema-agnostic. This is a fairly unique feature that allows users to “flip” schemas on the fly.

For example, it’s possible to present entire data sets to the engine based on an all VARCHAR schema (ie: make every column VARCHAR). Maybe you don’t know the real schema at the time, or perhaps you don’t care about it or perhaps the optimal schema can only be determined after some analysis is performed. Or perhaps there are inconsistencies in the data or DQ issues preventing a valid “load” based on a rigid schema. Or maybe it was just easier and quicker to export all the data as string types in the CSV. In either case, the XSPRADA engine will happily ingest that data. Later on, you can CAST each field as needed into a new table on the fly and run queries against the new model, or try others as needed. Similarly, ingestion validation is kept to a minimum by design. For example, it’s quite possible to load an 80-char string into a CHAR(3) field. This is not possible with conventional databases. The implications of this from a performance and flexibility angle are impressive. The XSPRADA database lends itself to internal transformation; hence it favors an ELT model, minus the “L”.

In recent years, there is increasing pressure to push the work of transformation into the DBMS, to enable parallel execution via SQL transformation scripts. This approach has been dubbed ELT since transformation is done after loading.

And finally, to support depth, an engine should allow rich analytics, provide an ability to “focus” in and out on the data, and provide a holistic un-segmented view of the entire data set:

“Modern data analyses involve increasingly sophisticated statistical methods… analysts often need to see both the forest and the trees in running these algorithms… The modern data warehouse should serve both as a deep data repository and as a sophisticated algorithmic runtime engine.”

A salient feature of the XSPRADA engine is its ability to handle multiple “types” of BI work at the same time. For example, it’s possible to mix OLAP, data mining, reporting and ad-hoc workloads simultaneously on the same data (and all of it) without resorting to “optimization” tricks for each mode. Similarly, the need for logical database partitioning doesn’t exist in the XSPRADA engine. Duplicating and re-modeling data islands on separate databases (physical or logical) for use by different departments is neither necessary nor recommended.

In an OLAP use case, there is no need to pre-define or load multidimensional cubes. The very act of querying consistently (meaning more than once) based on fact and dimension axes causes the engine to realize that this particular section of data is being accessed “multi-dimensionally”. It then starts cubing information internally, aggregating as indicated (if needed) by incoming queries. In this mode, perhaps the engine will decide a columnar storage approach is optimal and will re-structure the data accordingly. In a data mining use case, the approach is likely different because incoming queries are “incremental” (often pinpointed) and results are used to generate new queries without pre-determined patterns. The engine will likely start by eliminating vast “wasteland” areas of the data (the forest) from consideration as needed, then proceed to optimize specific islands of interest (the trees) as they become more relevant in the queries.

So overall, I think the XSPRADA analytical engine was indeed designed with “MAD-ness” from the get-go, even if the term didn’t exist years ago. It’s the approach and the philosophy that really matters. In that respect, we’re definitely headed for the MAD-house :)