"Koalas" provides the #Pandas #dataframes API on top of Apache #SPARK for #exploratory-data-analysis with better #performance #toread
on 02024-12-12"Dask" is another distributed #dataframes system that “uses #Pandas under the hood” for #exploratory-data-analysis with better #performance #toread
on 02024-12-12"Modin" scales #Pandas notebooks across clusters for #exploratory-data-analysis with better #performance #toread #dataframes
on 02024-12-12#compilers for #Pandas to get 1000× better #performance on #exploratory-data-analysis: "Dias" by #Baziotis
on 02024-12-12Wes McKinney started writing #Pandas in 02008 and subsequently worked on something called Badger, and then #Ibis, Arrow, Feather, and #Parquet. Here he explains the unifying thread, but this is mostly about #Arrow.
on 02024-11-22"Polars" is a #dataframes system similar to #Pandas but with up to 30× faster performance and IIRC a nicer API, available in not just Python but also Rust and node.js
on 02024-11-22#Pandas has a method for computing pairwise correlations of all columns using Pearson, Spearman, or Kendall. #statistics
on 02024-11-21"Visidata" is a spreadsheety GPL3 data visualization #TUI program that looks pretty impressive. #infovis in #Python, supporting #TSV, #JSON, #CSV, #SQLite, #Pandas, and #HDF5
on 02024-09-03Getting the time of day out of timestamp/datetime columns in #Pandas #databases is df['date_col'].dt.time. #dates
“Seaborn is a library for making attractive and informative #statistics #graphics in #Python. It is built on top of matplotlib and tightly integrated with the PyData stack, including support for #numpy and #pandas data structures and statistical routines from scipy and statsmodels.”
on 02015-11-16#pandas #Python vs. #R. includes things like #random-forest modeling, linear regression, and web scraping #machinelearning
on 02015-10-14How to upsample and downsample with the .resample method in #Pandas
Being able to handle datasets bigger than memory is the main reason a lot of people are still using SAS when they would otherwise prefer #Pandas or #R.
on 02015-08-15#Pandas #time-series downsampling uses the how parameter to aggregate points, while upsampling uses the fill_method parameter, generally with "pad".
#Guix now has #Pandas, as of 2015-05-15, thanks to Richard Wurmus.
on 02015-08-11my thoughts on a principled #APL, which is awfully similar to #Pandas, although I didn’t notice that at the time!
on 02015-08-05Wes McKinney’s #pandas tour #IPython notebook from his 10-minute Pandas demo video.
on 02015-08-05how to do dates in #pandas
on 02015-08-05Timedeltas in #Pandas.
on 02015-08-05How to do a relational join on #Pandas DataFrame objects.
on 02015-08-05#Pandas provides an efficient #EMA for #numpy.
on 02015-08-05