this post was submitted on 11 Aug 2024
-43 points (12.3% liked)

Data is Beautiful

5020 readers
8 users here now

A place to share and discuss visual representations of data: Graphs, charts, maps, etc.

DataIsBeautiful is for visualizations that effectively convey information. Aesthetics are an important part of information visualization, but pretty pictures are not the sole aim of this subreddit.

A place to share and discuss visual representations of data: Graphs, charts, maps, etc.

  A post must be (or contain) a qualifying data visualization.

  Directly link to the original source article of the visualization
    Original source article doesn't mean the original source image. Link to the full page of the source article as a link-type submission.
    If you made the visualization yourself, tag it as [OC]

  [OC] posts must state the data source(s) and tool(s) used in the first top-level comment on their submission.

  DO NOT claim "[OC]" for diagrams that are not yours.

  All diagrams must have at least one computer generated element.

  No reposts of popular posts within 1 month.

  Post titles must describe the data plainly without using sensationalized headlines. Clickbait posts will be removed.

  Posts involving American Politics, or contentious topics in American media, are permissible only on Thursdays (ET).

  Posts involving Personal Data are permissible only on Mondays (ET).

Please read through our FAQ if you are new to posting on DataIsBeautiful. Commenting Rules

Don't be intentionally rude, ever.

Comments should be constructive and related to the visual presented. Special attention is given to root-level comments.

Short comments and low effort replies are automatically removed.

Hate Speech and dogwhistling are not tolerated and will result in an immediate ban.

Personal attacks and rabble-rousing will be removed.

Moderators reserve discretion when issuing bans for inappropriate comments. Bans are also subject to you forfeiting all of your comments in this community.

Originally r/DataisBeautiful

founded 2 years ago
MODERATORS
 

▶️ Total olympic medals won in Paris 2024 and Human Development Index 🏅

@dataisbeautiful

➡️ https://www.businesstimes.com.sg/opinion-features/what-olympic-medal-table-really-tells-us

After reading the article we made this #boxplot using #LabPlot, an open source data analysis and visualization software.

The plot doesn't provide answers, it rather invites some thinking.

#Olympics #Olympics2024 #France #China #USA #UnitedStates #UnitedKingdom #UK #Brazil #Australia #Japan #Italy #Canada #Germany #Italy #Netherlands #DataAnalysis #DataScience #OpenSource #FOSS

top 21 comments
sorted by: hot top controversial new old
[–] [email protected] 15 points 4 months ago* (last edited 4 months ago)

This is not beautiful, this is confusing.
Why are the Nederlands with a index of 10 more ro the right than Australia that happens to have an index of 10 as well?
The "information" of the x-axis is completely random or so it seems.
If you plotted medals (y-axis) over index (x-axis) there might be information in there.
Bit this? C'mon, thats embarrassing.

[–] clutchtwopointzero 9 points 4 months ago (1 children)

"doesn't provides answers but invites thinking"... Nope. Doesn't even help that as the X-Axis is unlabelled

[–] [email protected] -5 points 4 months ago* (last edited 4 months ago) (2 children)

@clutchtwopointzero

A boxplot is a 1-dimensional plot. The data points are jittered along the x-axis to make them less crowded.

More on boxplots here:

➡️ https://labplot.kde.org/2021/08/11/box-plot/
➡️ https://userbase.kde.org/LabPlot/2DPlotting/BoxPlot

[–] clutchtwopointzero 5 points 4 months ago* (last edited 4 months ago)

Yeah, not a good way to visualize as the relationship between medal count and HDI is not obvious as only outliers get highlighted and the lack of information on other countries actually invite doubt as to the story that the plot is trying to tell (for example, Singapore and Hong Kong have extremely high HDI but the sheer smallness of their population is a factor against a higher medal count). There's nothing wrong with a traditional 2D scatter plot and axis-related box plots plotted against each axis separately

[–] WhatAmLemmy 2 points 4 months ago

This viz is shit because nobody can understand it.

1 dimensional visualizations are fucking stupid. The whole point of data viz is to simplify complex information; not make simple info complex.

[–] [email protected] 6 points 4 months ago (1 children)

A boxplot is a visualization tool to quickly get an idea of how the data is distributed. In this population the outliers are so large that the info the real box + whiskers give is very low.

In your title you suggest investigating a relationship between total Olympic medals and HDI - why not choose a scatter plot here?

That the number in square brackets refers to the HDI rank only get's clear on the second look.

The outliers being distributed over the X-Axis is confusing.

Sorry but this visualization is not beautiful, rather the wrong method used that cannot display the hypothesis stated in the title.

[–] clutchtwopointzero 3 points 4 months ago* (last edited 4 months ago)

The choice of only highlighting the HDI of the outliers makes one wonder what the rest of the data is hiding and whether this graph is hiding the truth from the data to tell a biased narrative.

Also, box plots only work on a single dimension.

[–] [email protected] 4 points 4 months ago (1 children)

@LabPlot @dataisbeautiful
Doesn’t make sense unless you calculate in population size. Best way to do this is to have “# medals per capita ratio” on the vertical axis instead of simply # medals.

[–] stupidcasey 2 points 4 months ago (1 children)

This doesn’t make any sense at all, it’s trying to force correlation to be causation as some political agenda that I can’t quite understand.

[–] [email protected] 0 points 4 months ago* (last edited 4 months ago) (2 children)

@stupidcasey Ok, let me explain: if you look at the chart it looks like the US is doing much much better than Australia. Twice the # of medals and about same score on human development index. Truth is US has over 12x the population of Australia.
If you adjust per my suggestion you’d see that Australia is doing ~6x better than US instead of US doing ~2x better than Australia as it is in the chart now. Much more realistic, isn’t it?

[–] stupidcasey 1 points 4 months ago (1 children)

IM hearing a lot of correlation and not A whole lot of causation there, did the US have 12x the people competing in the Olympics?, did Australia pick its people from an even distribution of its populous or maybe just maybe did they Cherry pick from places that are better than the US like Sydney or Melbourne?

[–] [email protected] 1 points 4 months ago (1 children)

@stupidcasey
if a country has 12x a pool of people to pick their best athletes from, wouldn’t you agree that would hugely increase their winning chances?
If two schools compete in a chess match, 1 school has 100 students, the other 1200 students, and they both send their best chess player, with all other factors being equal, who would you put your money on?

[–] stupidcasey 2 points 4 months ago (1 children)

Nope, I think both countries have more people than it would be possible to evaluate, also that’s population not living standards, also training has more to do with it than the individual initially picked also the amount of money it takes to train an athlete is such a small percentage of either country’s GDP that money just doesn’t matter either,

All that together plus the plethora of other variables makes this correlation not causation.

[–] [email protected] 1 points 4 months ago (1 children)

@stupidcasey @stupidcasey so you really seriously think based on that above chart that US and then China are the ‘best performing’ countries and the fact that they have a huge population has nothing to do with it????

[–] stupidcasey 1 points 4 months ago (1 children)

No I don’t, that chart is the human development index. Why would I draw a conclusion about population from a chart about human development index? They have nothing in common.

[–] [email protected] 1 points 4 months ago (1 children)

@stupidcasey I like have conversations, but don’t appreciate your tone of voice. That’s why I’ll block you. This is not Twitter.

[–] stupidcasey 1 points 4 months ago

M k, by👋👋👋

[–] [email protected] 0 points 4 months ago

@stupidcasey And what all of this has to do with any political agenda; beats me! 🤣🤣🤣

[–] [email protected] 2 points 4 months ago* (last edited 4 months ago)

@dataisbeautiful

Thank you for all your comments. A jittering of data points along the x-axis was used to avoid over-plotting. But yes, a scatter plot with a boxplot attached along the y-axis (to show outliers) may be more informative in this case.

[–] [email protected] 2 points 4 months ago

OP has tagged Canada but it’s not shown in the plot.

[–] [email protected] 2 points 4 months ago

Please for gods sake use per capita