Tyler Vigen’s list of spurious correlations, such as the very tight correlation between bachelor’s degrees awarded in psychology and the number of groundskeepers in Utah (R² = 0.980). #humor #statistics #epistemology
on 02026-08-18#USA #manufacturing #statistics on #batteries: skyrocketing to 240 (240 whats?) since 02020, when it had dipped to 65. It turns out that this is 240% of the index year of 02017, seasonally adjusted.
on 02026-06-15longer explanation of #Kullback-Leibler divergence. #math #statistics #toread
on 02026-04-08pretty explanation of #Kullback-Leibler divergence. #math #statistics #toread
on 02026-04-08#Veritasium #video about #power-laws like the "St. Petersburg Paradox" in #statistics, where the expectation of a heavy-tailed distribution diverges. Sometimes “two exponentials are dancing together to make a power law”, moving on to earthquake magnitudes and how the Ising model of ferromagnetism at the Curie temperature is scale-free, and then how a simple model of forest fires produces self-organized criticality. Also, in Per Bak’s sandpile automata. And how VC firms, wildfire insurance, streaming video platforms, and book publishing are governed by power-law distributions. And Albert-László Barabási and Réka Albert’s simulation of power-law link behavior on the internet.
on 02026-01-03#OECD #statistics on #infant-mortality. Argentina is 8.2 deaths per 1000 live births, Turkey is 9.8, Mexico is 13.4, Chile is 6.5, USA and Slovakia are 5.6 #medicine
on 02025-11-05Electrical injuries kill 1000 people a year in #USA. #statistics #safety #paper
on 02025-10-05Gelman’s #statistics blog
on 02025-09-28Gelman’s #statistics #ebook “Bayesian Data Analysis, Third edition” #PDF #toread (677 pp.)
on 02025-09-28#statistics on #USA drug overdose deaths, which reached an age-adjusted peak of 32.3 per 100k population per year in 02022, then declines slightly in 02023.
on 02025-04-05Jeremy Giles reports in #Foreign-Policy that #Bukele’s #homicide #statistics are fake because they stopped counting (at least) mass graves, police killings, and killings in prison as “homicide”: “Since the start of 2021, at least 171 unmarked graves have been discovered (68 in 2021, 43 in 2022, and 60 in 2023) and entirely left out of the nation’s homicide figures. (...) 1.6 percent of the Salvadoran population is now behind bars—more than twice the incarceration rate of Cuba and Rwanda and more than four times that of the United States.” Includes month-by-month homicide plots! #El-Salvador #politics #crime
on 02024-12-21#statistics in Python, including specifically multivariate linear regression, including scikit-learn and statsmodels OLS.
on 02024-12-19#SciPy can do univariate linear regression with scipy.stats.linregress #statistics
on 02024-12-19#news #video with #economics #statistics about #Argentina: +3.9% GDP in the third quarter of 02024, -1.7% the previous one, -2.1% the one before that, and -1.9% from the last quarter of 02023. That still leaves us 2% below the last quarter of 02023. The talking head in the video, however, says it leaves us 0.1% below. It seems like this is a combination of calculation errors (maybe he’s naïvely adding percentages instead of multiplying, or maybe I’m using incorrectly rounded numbers) and not counting the -1.9%. He also says that the quarter-on-quarter growth in construction activity was 6.9% and in industrial production was 13.6%. The year-on-year figures he shows seem to show a dismaying backslide towards a resource-extraction economy: 32.0% growth in agriculture, 7.3% growth in mining, 7% growth in fishing, 2.9% growth in hospitality [...] and a 10.7% drop in manufacturing and a 14.9% drop in construction. Which I guess are the sectors that rebounded most sharply in this last quarter. And apparently last year there was a drought.
on 02024-12-17#Australia #energy #statistics: 18 186 063 TJ/year (576GW)
on 02024-12-09#Pandas has a method for computing pairwise correlations of all columns using Pearson, Spearman, or Kendall. #statistics
on 02024-11-21#statistics on the inflation-adjusted interest rate on #USA treasury bonds at different maturities and different dates. #economics #data #finance
on 02024-11-03#statistics on the one-year inflation-adjusted interest rate #USA #finance #economics #data
on 02024-11-03#Economics #statistics for #Argentina, like how our interest rate is 40%
on 02024-08-20#nuclear #energy #safety and #statistics
on 02023-07-01#math #statistics proof by Thomas Royen of the Gaussian correlation inequality; nearly ignored because he wrote it in Word and published it in “the Far East Journal of Theoretical Statistics, a periodical based in Allahabad, India...which...listed Royen as an editor”. Not a joke!
on 02017-05-02#Statistics #math #paper #toread on a very generic explanation for Zipf’s law
on 02017-03-11Brad DeLong uses Michael Kremer’s #statistics on the #history of human #population to find robustly that the British #Industrial-Revolution departed from the historical population growth trend.
on 02016-08-01Six basic points from the American Statistical Association on p-value #statistics: “① P-values can indicate how incompatible the data are with a specified statistical model. ② P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone. ③ Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. ④ Proper inference requires full reporting and transparency. ⑤ A p-value, or statistical significance, does not measure the size of an effect or the importance of a result. ⑥ By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.”
on 02016-03-08“Seaborn is a library for making attractive and informative #statistics #graphics in #Python. It is built on top of matplotlib and tightly integrated with the PyData stack, including support for #numpy and #pandas data structures and statistical routines from scipy and statsmodels.”
on 02015-11-16"TensorFlow" is a high-throughput #dataflow array computing library and IDE with built-in reverse-mode #automatic-differentiation for optimization, with Python and C++ APIs and #IPython integration. At this moment in history, the growth of computer power has made a bunch of important #DSP and statistical tasks just feasible, so we are seeing things like self-driving cars, superhuman image recognition, and so on. But it’s been very difficult to take advantage of the available computational power, because it’s in the form of GPUs and clusters. So this is designed to make it easy to do exactly these things, and to scale them with your available computing power, along with libraries of the latest tricks in neural networks, machine learning (which is pretty close to "statistics").
on 02015-11-09“P-values are not error probabilities”, with an overview of the #history of #statistics with Fisher’s evidential p value contrasted with the Neyman–Pearson α.
on 02015-08-13#Probabilistic-programming allows you to condition the execution of a nondeterministic program on observations, reasoning backwards from the observations to the conditional distribution of the original stochastic variables, thus allowing you to express #statistics models in a Turing-complete form. (However, some PPS languages are not Turing-complete.) Well-known languages are Venture, Figaro, and Berkeley’s BLOG, but there are also Church, Haraku, Chimple/Dimple, Stan, and Infer.net. #toread
on 02015-08-13A psychology textbook #ebook chapter about #nonparametric #statistics which omits even mentioning Kolmogorov-Smirnov Apparently, that’s because this is purely about nonparametric statistics on ordinal data, and the K-S test is really only applicable to data on at least an interval scale of measurement.
on 02015-08-13Another #nonparametric #statistics test, like Wilcoxon and Kolmogorov-Smirnov. “Nearly as efficient as the t-test on normal distributions.” That’s very exciting!
on 02015-08-10is the other main #nonparametric #statistics hypothesis test for comparing two distributions. I'm pretty sure it doesn't make sense on ordinal measurements, since the CDF it measures will always be the straight line y=x/n for sufficiently-precisely-measured ordinal data, regardless of the underlying distribution. By contrast, Wilcoxon and Mann–Whitney are applicable to ordinal variables.
on 02015-08-10Wilcoxon (wilcox.text in R, scipy.stats.wilcoxon in SciPy) is a #nonparametric #statistics hypothesis test.
on 02015-08-10#bayesian #statistics #philosophy #toread
on 02015-08-05Commentary on Breiman's #statistics vs. #machine-learning (more or less) paper
on 02015-08-05#Statistics vs. #machine-learning: the same thing really
on 02015-08-05#statistics and algorithmic vs. data modeling (Leo Breiman is firmly in the camp in favor of algorithmic modeling rather than data modeling)
on 02015-08-05Pitfalls of #statistics.
on 02015-08-05