{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "@id": "https://www.vidyasource.com/blog/java-is-dysfunctional-with-big-data/",
  "url": "https://www.vidyasource.com/blog/java-is-dysfunctional-with-big-data/",
  "mainEntityOfPage": "https://www.vidyasource.com/blog/java-is-dysfunctional-with-big-data/",
  "headline": "Java is Dysfunctional with Big Data",
  "description": "Java is great, but there are far better options for big data analytics.",
  "datePublished": "2013-10-27T00:00:00.000Z",
  "author": {
    "@type": "Person",
    "name": "Neil Chaudhuri",
    "jobTitle": "President",
    "url": "https://www.linkedin.com/in/neil-chaudhuri/",
    "sameAs": "https://www.linkedin.com/in/neil-chaudhuri/"
  },
  "publisher": {
    "@id": "https://www.vidyasource.com/#organization"
  },
  "image": "https://www.vidyasource.com/img/blog/big-data.jpg",
  "keywords": [
    "Java",
    "Scala",
    "Hadoop",
    "Apache Spark",
    "Clojure",
    "Cascalog",
    "MapReduce",
    "Programming",
    "Functional Programming",
    "Big Data",
    "Analytics"
  ],
  "articleBody": "Let me first say I love Java. There is a reason it’s the most popular programming language in the world. For me\npersonally, I made a career out of building systems in Java, and I even teach a course in Java.\n\nBut when it comes to Big Data, Java simply doesn’t cut it.\n\nEverybody knows functional languages have enjoyed a renaissance as Big Data has become a thing. And for good reason:\nimmutable (or mostly immutable) state, lazy evaluation, the natural fit with recursion, and so on. There are a lot of\nresources out there like [this](http://cafe.elharo.com/programming/java-programming/why-functional-programming-in-java-is-dangerous/)\nblog post where you can explore those concepts and the passions they arouse.\n\nBut for all the computer science theory, inelegant blogger arguments, and angry feedback from purists still bitter they\nnever went to prom, most of us have jobs to do, analytics to run, and customers to please. We need tools that allow us\nto get the most done the most quickly. That just isn’t Java.\n\nAlmost everyone who has used the Hadoop stack learned MapReduce by writing the canonical\n[Word Count](http://developer.yahoo.com/hadoop/tutorial/module4.html#wordcount) example in Java and did just fine, so\nwhat’s the problem? We have evolved past low-level Hadoop programming and operate at a much higher level of abstraction.\nBig Data analytics is really a series of operations on collections--initialization followed by some combination of\nfiltering operations (extracting only elements that satisfy some condition), mapping operations (transforming elements),\nand reducing operations (aggregating elements) that produce new collections.\n\nJava is terrible at collections. You don’t think so? Let’s take a simple scenario where you have a collection of numbers\nrepresenting your data. I want to do the following:\n\n1. Filter out all even numbers\n2. Map the result by multiplying each number by 3\n3. Reduce the result to their sum\n\nPretty simple. For example [1,2,3,4,5] --> [2,4] -> [6, 12] -> 18\n\nHere is a reasonable implementation of this algorithm in Java.\n\n~~~java \npublic class CollectionsDemo {\n    public CollectionsDemo() {\n        List<Integer> data = BigDataGenerator.data();\n        int finalResult = sum(map(filter(data)));\n    }\n\n    public List<Integer> filter(List<Integer> data) {\n        List<Integer> evens = new ArrayList<Integer>();\n        for (Integer number : data) {\n            if (number % 2 == 0) {\n                evens.add(number);\n            }\n        }\n\n        return evens;\n    }\n\n    public List<Integer> map(List<Integer> data) {\n        List<Integer> mappedData = new ArrayList<Integer>();\n        for (Integer number : data) {\n            mappedData.add(3 * number);\n        }\n\n        return mappedData;\n    }\n\n    public int sum(List<Integer> data) {\n        int sum = 0;\n        for (Integer number : data) {\n            sum += number;\n        }\n\n        return sum;\n    }\n\n    public static void main(String[] args) {\n        new CollectionsDemo();\n\n    }\n}\n~~~\n\nThat’s a lot of code to perform three trivial operations, and it’s because of all the boilerplate. Sometimes we forget\nhow much boilerplate there is in Java because our IDE’s generate it for us.\n\nIf we are going to be productive with Big Data, we need languages that allow us to manipulate collections much more easily.\nJustin Timberlake brought [sexy back](http://www.youtube.com/watch?v=3gOHvDP_vCs); Big Data brought functional languages back.\n\nI mentioned all the technical details like immutability that make functional languages well-suited to Big Data, but it’s\ntheir suitability to Big Data *developers* that really matters. Developers need a compelling reason to make the effort\nto learn functional languages, which are much more difficult to understand than an object-oriented language like Java.\n\nTo prove my point, here is the same algorithm in Scala, another language for the JVM that combines object-oriented and\nfunctional paradigms:\n\n~~~scala\nval finalResult = data.filter(_ % 2 == 0).map(3 * _).sum\n~~~\n\nI know what you’re thinking. [Really?](http://www.thedailybeast.com/videos/2013/05/19/really-amy-poehler-returns-to-snl.html)\nYup. That’s it. And it’s readable too.\n\nAnd here is the same algorithm in Clojure, a Lisp (so a functional language) for the JVM and other environments:\n\n~~~clojure\n(reduce +\n  (map #(* 3 %)\n  (filter even? (data))))\n~~~\n\nThis might be slightly less readable but still worth it to learn.\n\nI mention Scala and Clojure because they are the languages of choice for two preeminent Big Data tools.\n\n[Apache Spark](http://spark.incubator.apache.org/) was originally created by [AMPLab](https://amplab.cs.berkeley.edu/)\nat UC Berkeley and is now in incubation at Apache. Spark utilizes cluster memory to minimize the disk reads and writes\nthat slow down MapReduce. Spark is written in Scala and exposes a Scala API.\n\n[Cascalog](https://github.com/nathanmarz/cascalog) is a tool created by data overlord [Nathan Marz](https://twitter.com/nathanmarz),\nformerly of Twitter and quite possibly the world’s foremost expert on Big Data batch and streaming analytics. Cascalog\nis an abstraction written in Clojure on top of the core Hadoop stack (MapReduce, Hive, Pig) that helps you avoid\nHadoop’s low level details.\n\nRemember what a pain [Word Count](http://developer.yahoo.com/hadoop/tutorial/module4.html#wordcount)  is in Java? Look\nat the difference with these alternatives.\n\nWord Count in Spark:\n\n~~~scala\nfile.flatMap(line => line.split(\" \"))\n    .map(word => (word, 1))\n    .reduceByKey(_ + _)\n~~~\n\nWord Count in Cascalog:\n\n~~~clojure\n(?<- (stdout)\n     [?word ?count]\n     (sentence ?line)\n     (tokenise ?line :> ?word)\n     (c/count ?count))\n~~~\n\nIs there any doubt that Java is in over its head with Big Data? A developer has to be pretty [lazy](http://www.youtube.com/watch?v=Px5TWbc4xQo)\nnot to make the effort to learn at least one functional language.\n\nBut if you are, all is not lost. Both Spark and Cascalog (as [JCascalog](https://github.com/nathanmarz/cascalog/wiki/JCascalog))\nhava Java API’s too. As cumbersome as Java is, they are realistic enough to know that minimizing the barrier to entry for developers *and* their organizations is important.\n\nAnd there are efforts underway to make Java itself functional. Right now, you can use anonymous inner classes to fake it,\nor you can look into third-party solutions like Google Guava’s [Iterables](http://docs.guava-libraries.googlecode.com/git/javadoc/com/google/common/collect/Iterables.html) or random [frameworks](https://github.com/functionaljava/functionaljava) of dubious longevity. Most promising is the possibility that the language itself will evolve--that Java 8 will represent...finally...the arrival of [functional Java](http://blog.agiledeveloper.com/2013/01/functional-programming-in-java-is-quite.html).\n\nBut keep in mind two things. First, they’ve been promising functional programming in Java since Java 6. I’ve been promised\na [Justice League](http://media.dcentertainment.com/sites/default/files/character_bio_576_justiceleague.jpg) movie too. Second,\neven if Java 8 comes out functional tomorrow, it isn’t like organizations will drop everything to make the switch. Project\nand technical leads will assess the risks and rewards of moving to Java 8. It could be years before the switch happens on your project.\n\nSo if you’re a Java developer who wants to get good at Big Data, start learning Scala or Clojure now. Python wouldn’t\nhurt either. Why should Commander Data be the only one who is [fully functional](http://www.youtube.com/watch?v=9ev1ec0Z0GI)?"
}