Embarrassing how bad the #performance of #big-data things like Elastic Map Reduce and #Hadoop can be when used on only a couple of gigabytes of data; the Configuration to Outperform a Single Thread is apparently a lot bigger than that because this guy got the answer from 1.75 gigabytes in 12 seconds on his laptop (in parallel, using xargs -P4) instead of the 26 minutes the Hadoop cluster needed. He gets slightly better performance using mawk instead of gawk.
Impala is a #database built on #Spark and #Hadoop (and Hive) that gets truly impressive #SQL speeds; here’s how to install it
on 02015-08-05