{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "@id": "https://www.vidyasource.com/blog/lighting-spark-hbase/",
  "url": "https://www.vidyasource.com/blog/lighting-spark-hbase/",
  "mainEntityOfPage": "https://www.vidyasource.com/blog/lighting-spark-hbase/",
  "headline": "Lighting a Spark with HBase",
  "description": "Apache Spark is great for Hadoop analytics, and it works just fine with HBase.",
  "datePublished": "2014-01-25T00:00:00.000Z",
  "author": {
    "@type": "Person",
    "name": "Neil Chaudhuri",
    "jobTitle": "President",
    "url": "https://www.linkedin.com/in/neil-chaudhuri/",
    "sameAs": "https://www.linkedin.com/in/neil-chaudhuri/"
  },
  "publisher": {
    "@id": "https://www.vidyasource.com/#organization"
  },
  "image": "https://www.vidyasource.com/img/blog/spark.png",
  "keywords": [
    "Java",
    "Scala",
    "Hadoop",
    "Apache Spark",
    "HBase",
    "MapReduce",
    "Hive",
    "HDFS",
    "Big Data",
    "Programming",
    "Functional Programming",
    "Open Source"
  ],
  "articleBody": "Although most developers and users are still feeling their way through [Hadoop](http://hadoop.apache.org/) and (more\nspecifically [MapReduce](https://hadoop.apache.org/docs/r1.2.1/mapred_tutorial.html)), the truth is Google wrote\nthat [paper](http://static.googleusercontent.com/media/research.google.com/en/us/archive/mapreduce-osdi04.pdf) in 2004.\nThat’s ten years ago! *[Million Dollar Baby](http://www.imdb.com/title/tt0405159/)* won Best Picture that year.\n*[Yeah!](http://www.youtube.com/watch?v=eSPhCS-15eE)* by Usher, Lil John and Ludacris topped the charts in the United\nStates. And Facebook had only just started to kill work productivity and violate\nyour privacy.\n\nAs long ago as that feels, it is an eternity in technology. Google has of course moved on way past MapReduce to things\nlike [Dremel](http://static.googleusercontent.com/media/research.google.com/en/us/pubs/archive/36632.pdf). The rest of\nus aren’t moving so quickly, but that doesn’t mean we still need to plod along writing low-level MapReduce code.\n\nSeveral abstractions have come along to make cloud-scale analytics easier. One of them is\n[Apache Spark](http://spark.incubator.apache.org/). Spark was originally created by [AMPLab](https://amplab.cs.berkeley.edu/)\nat UC Berkeley and is now in incubation at Apache. It utilizes cluster memory to minimize the disk reads and writes that\nslow down MapReduce. Spark is written in Scala and exposes both a Scala and Java\nAPI. Though [as I’ve said before](https://www.vidyasource.com/blog/java-is-dysfunctional-with-big-data/),\nonce you get past the learning curve, I don’t see how anyone would prefer Java to Scala for analytics.\n\nPlease take a look at the [Spark documentation](http://spark.incubator.apache.org/docs/latest/scala-programming-guide.html)\nto learn about the [*RDD*](http://spark.incubator.apache.org/docs/latest/api/core/index.html#org.apache.spark.RDD.RDD)\nabstraction. *RDD*’s are basically fancy arrays. You load data into them and perform whatever combination of\noperations--maps, filters, reducers, *etc.*-- you want. *RDD*’s make analytics so much more fun to write than canonical\nMapReduce. In the end, Spark doesn't just run faster; it lets us write faster.\n\nThe web has a bunch of examples of using Spark with Hadoop components like\n[HDFS](http://hadoop.apache.org/docs/stable1/hdfs_design.html) and [Hive](https://cwiki.apache.org/confluence/display/Hive/GettingStarted)\n(via [Shark](https://github.com/amplab/shark/wiki), also made by AMPLab), but there is surprisingly little on using\nSpark to create *RDD*’s from HBase, the Hadoop database. If you don’t know HBase, check out this excellent\n[presentation](http://www.slideshare.net/cloudera/5-h-base-schemahbasecon2012) by\n[Ian Varley](https://twitter.com/thefutureian). It’s just a cloud-scale key-value store.\n\nThe Spark [Quick Start documentation](http://spark.incubator.apache.org/docs/latest/quick-start.html) says, “*RDD*s can\nbe created from Hadoop InputFormats.” You may know that\n[InputFormat](http://hadoop.apache.org/docs/current/api/org/apache/hadoop/mapred/InputFormat.html) is the Hadoop\nabstraction for anything that can be processed in a MapReduce job. As it turns out, HBase uses a\n[TableInputFormat](http://hbase.apache.org/apidocs/org/apache/hadoop/hbase/mapreduce/TableInputFormat.html), so it\nshould be possible to use Spark with HBase.\n\nIt turns out that it is.\n\nAs the [Scaladoc](http://spark.incubator.apache.org/docs/latest/api/core/index.html#org.apache.spark.RDD.RDD) for *RDD*\nshows, there are numerous concrete `RDD` implementations--each best suited for different situations. The\n[NewHadoopRDD](http://spark.incubator.apache.org/docs/latest/api/core/index.html#org.apache.spark.RDD.NewHadoopRDD)\nis great for reading data stored in later versions of Hadoop, which is exactly what we need here.\n\nCheck out this code.\n\n~~~scala\nval sparkContext = new SparkContext(\"local\", \"Simple App\")\nval hbaseConfiguration = (hbaseConfigFileName: String, tableName: String) => {\n  val hbaseConfiguration = HBaseConfiguration.create()\n  hbaseConfiguration.addResource(hbaseConfigFileName)\n  hbaseConfiguration.set(TableInputFormat.INPUT_TABLE, tableName)\n  hbaseConfiguration\n  }\nval rdd = sparkContext.newAPIHadoopRDD(\n  hbaseConfiguration(\"/path/to/hbase-site.xml\", \"table-with-data\"),\n  classOf[TableInputFormat],\n  classOf[ImmutableBytesWritable],\n  classOf[Result]\n)\nimport scala.collection.JavaConverters._\nrdd\n  .map(tuple => tuple._2)\n  .map(result => result.getColumn(\"columnFamily\".getBytes(), \"columnQualifier\".getBytes()))\n  .map(keyValues => {\n  keyValues.asScala.reduceLeft {\n    (a, b) => if (a.getTimestamp > b.getTimestamp) a else b\n  }.getValue\n})\n~~~\n\nOnce we instantiate the [SparkContext](http://spark.incubator.apache.org/docs/latest/api/core/index.html#org.apache.spark.SparkContext)\nfor the local machine, we write an anonymous function to create an\n[HBaseConfiguration](http://hbase.apache.org/apidocs/org/apache/hadoop/hbase/HBaseConfiguration.html), which will enable\nus to add the HBase configuration files to any Hadoop ones we have as well as the name of the table\nwe want to read. That’s an HBase thing--not a Spark thing.\n\nThen we create an instance of `NewHadoopRDD` with the `SparkContext` instance and three Java\n[Class](http://docs.oracle.com/javase/7/docs/api/java/lang/Class.html) objects related to HBase:\n\n* One for the `InputFormat`, `TableInputFormat`, as mentioned earlier\n* One for the key type, which on a table scan is always [ImmutableBytesWritable](http://hbase.apache.org/apidocs/org/apache/hadoop/hbase/io/ImmutableBytesWritable.html)\n* One for the value type, which on a table scan is always [Result](http://hbase.apache.org/apidocs/org/apache/hadoop/hbase/client/Result.html)\n\nThen we call our anonymous function with the location of [hbase-site.xml](http://hbase.apache.org/book/config.files.html)\nand the name of the table we want to read, `table-with-data`,\nto provide Spark with the necessary HBase configuration to construct the `RDD`. Once we have the `RDD`, we can perform all\nthe usual operations on HBase we usually see with the more conventional usage of `RDD`s in Spark.\n\nThe code example does feature some HBase-specific operations though. When the *RDD* is constructed, it loads all the\ndata in `table-with-data` as (`ImmutableBytesWritable`, `Result`) tuples--key-value pairs as you would expect. For each\none of these pairs, we grab the second item in the tuple (the `Result`) and\nget the column\nfrom it marked with a specific [column family and column qualifier](http://hbase.apache.org/book/columnfamily.html).\n\n[As advocated](http://tek-tips.nethawk.net/new-paradigm-and-thinking-required-for-massively-distributed-and-complex-systems/)\nby big data overlord [Nathan Marz](https://twitter.com/nathanmarz), HBase has no notion of deletes; new values are just\nappended. Consequently, each column is really a collection, and after converting the Java collection returned by the HBase API\ninto a Scala collection, the last *map* call extracts the latest--and therefore\nthe “right”--[KeyValue](http://hbase.apache.org/apidocs/org/apache/hadoop/hbase/KeyValue.html). We now have a collection\nof all the current data in each row.\n\nIf that is unclear, let's summarize the mapping operations this way:\n\n* Load an `RDD` of (`ImmutableBytesWritable`, `Result`) tuples from the table\n* Transform previous into an `RDD` of *Result*'s\n* Transform previous into an `RDD` of collections of `KeyValue`s (where the latest one is what we want). Collection of collections? [Oh no!](http://www.youtube.com/watch?v=Xpc0s9FsA1Q)\n* Transform previous into an RDD of `byte[]` (binary representations of the current data)\n\nSo using Spark, we can read the contents of an entire HBase table and perform one transformation after another--maybe\nalong with some filters, aggregations, and other operations--to do [ALL THE ANALYTICS](http://www.tjkelly.com/wp/wp-content/uploads/hyperbole-clean-all-the-things.jpg).\n\nAstute or experienced readers might have noticed that `getColumn` call in the second `map` operation is actually\ndeprecated in favor of `getColumnCells`,\nwhich returns a collection of [Cell](http://hbase.apache.org/apidocs/org/apache/hadoop/hbase/Cell.html)s ordered by\ntimestamp. That’s the superior approach. I just used `getColumn` as an excuse to demonstrate Spark’s ability to apply\nanonymous functions within `RDD` operations in the next `map` call.\n\nThe next time some analytics are in order, don’t bother with old-school MapReduce. The higher level of abstraction over anything with an\n`InputFormat`, including HBase, makes Spark a great choice for Hadoop analytics."
}