Embarrassing how bad the #performance of #big-data things like Elastic Map Reduce and #Hadoop can be when used on only a couple of gigabytes of data; the Configuration to Outperform a Single Thread is apparently a lot bigger than that because this guy got the answer from 1.75 gigabytes in 12 seconds on his laptop (in parallel, using xargs -P4) instead of the 26 minutes the Hadoop cluster needed. He gets slightly better performance using mawk instead of gawk.
What is the Configuration to Outperform a Single Thread? Surprisingly often it’s ridiculous or even unbounded. #performance #Spark #big-data
on 02015-08-05