Skip to main content

Posts

Showing posts with the label big data

Offline matching and the catalog from Dell

A couple of weeks ago I saw an ad on my phone for the new Dell XPS laptop. Out of curiosity I clicked on it, and started configuring a couple of options just for fun. I read a lot of good reviews about the upcoming XPS laptop, how the design rivals that of Apple’s MacBook laptops, and how certain configurations come with Linux preinstalled. I played around with the options, and configured two beast laptops just for fun, then I left and have not thought about it anymore, since I am not in the market for a new laptop, being satisfied with my Surface 3 laptop and all. A couple of weeks later, I received a Dell catalog for the first time, addressed to me and not to the usual “Current Resident”, that has more details about the XPS laptops, gaming laptops, desktops, and peripherals that Dell sells. The catalog also included a 15% discount, which was a nice touch. I wondered how I got that catalog, even though I have not explicitly sign up for it, nor request one, nor provide any informat...

A paper a day keeps the doctor away: BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data

The latest advances in Big Data systems have made storing and computing over large amounts of data more tractable than in the past. Users' expectations for how long a query should take to complete have not on the other hand   changed, and remain independent of the amount of data that needs to be processed. The expectation mismatch of query run time causes user frustration when iteratively exploring large data sets in search of an insight. How can we alleviate that frustration? BlinkDB offers users a way to balance result accuracy with query execution time : the users can either get quantifiably approximate answers very quickly, or they can elect to wait for a longer period of time to get more accurate results. BlinkDB accomplishes this tradeoff through the magic of dynamic sample selection, and an adaptive optimization framework. The authors start with an illustrative example of computing the average session time for all users in New York. If the table that stores users...

Enterprise Data Workflows with Cascading, by Paco Nathan, O'Reilly Media

For people interested in developing Hadoop analytic applications there is a plethora of options. The options range from writing low-level, hand-tuned Java map-reduce code, to using a higher level language to manipulate the data such as Pig and Hive. There are pros and cons for each option. For the first, the code becomes complex for anything other than the canonical word-count example, and for the latter, to do anything meaningful, you almost always end up augmenting the higher level language with user-defined functions written in a different language to regain power and flexibility, causing maintenance nightmares. A happy medium in between is to use one of the data-flow libraries for Hadoop, of which Cascading is one. Since Cascading has been around for some time, the online documentation is relatively mature, and includes a gentle introduction to the library, with example source code, and a well written user's guide. However this does not obviate the need for a book that desc...

PiQL tech talk

Big Data has been gaining a lot of press lately, and NoSQL even more. A lot of new software development is moving to using NoSQL databases to alleviate the scaling pains of traditional RDBMS systems when the data size grows very large. The NoSQL databases usually expose a simple key value based API, where you can set, retrieve, and scan values based on the value of the keys you're interested in. The API is sufficient for most applications, but sometimes you want more than the simple retrieval API; you want SQL, where you can join keys, and have more complex filtering rules. Here is where PiQL plays a role. Last week I attended a talk about PiQL by Michael Armbrust. In the talk, Michael introduced PiQL: a query system that runs on top of NoSQL key-value stores, and allows developers to write SQL queries that execute efficiently against the simple key-value retrieval API. The system operates in two stages: the first is the static analysis stage, and the second is the execution s...