Rendered at 23:41:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
JHonaker 6 hours ago [-]
After seeing the claim, I knew it would work because of a few intuitive things:
1. Things that are the result of a very large number of small perturbations follow a Gaussian distribution.
1a. Word embeddings are the result of millions/billions/trillions of small perturbations (i.e. gradient descent)
2. As you increase the dimension, d, of a multivariate Gaussian distribution, the actual mass of the data is increasingly concentrated in the thin spherical shell. For MVN distributions with standardized marginal dimensions, this is on the sphere of radius \sqrt{d}.
In fact when I searched "Gaussian curse of dimensionality" I got this at the top result: [1] which shows this phenomena perfectly. Or this one with pretty interactive pictures [2]
Does the mapping depend on the chose set, or could you just map all 25k words and pick 10 points that form a circle? From the process, they seemingly replaced single words in the set that didn't fit, which makes it sound like the position of the remaining words didn't change much from changing the set.
jamie-simon 11 hours ago [-]
interesting. yeah, there's probs some optimization that can be done there.
mph1027 17 hours ago [-]
The best bar bets are the ones where you're technically not cheating, just abusing linear algebra.
ZeroDayDreamer 15 hours ago [-]
Using high-dimensional statistics for beer money, not academic papers" is the correct use of the degree. The papers are where you prove the circle exists. The bar is where you find out it exists for everything
mdritch 22 hours ago [-]
I recently did some analysis on how different style prompts like "be succinct" and "avoid mannered prose" affect LLM outputs and I also saw a ~circular relationship with a 2d dimensional reduction (MDS) of the similarity matrix. I wonder if it's some artifact of PCA/MDS + Gemma's output dist
ViscountPenguin 22 hours ago [-]
A fun step-up would be to see if you can find an infinity shape by repeating the process with a set of points which are maximally far away from all points but 1 on the existing circle.
ovin_dal 19 hours ago [-]
Betting on this with high-dimensional stats sounds like a classic case of overfitting your opponent. Just buy them a beer.
cold_boot 20 hours ago [-]
Tried something similar predicting poker outcomes, ended up buying everyone a round. The math wasn't in my favor.
heronbank 19 hours ago [-]
My stats professor would be proud I'm using high-dimensional statistics for beer money, not academic papers.
Eastmill 19 hours ago [-]
Used a similar trick once to predict who'd buy coffee. Didn't win a beer, but definitely earned some weird looks.
1. Things that are the result of a very large number of small perturbations follow a Gaussian distribution.
1a. Word embeddings are the result of millions/billions/trillions of small perturbations (i.e. gradient descent)
2. As you increase the dimension, d, of a multivariate Gaussian distribution, the actual mass of the data is increasingly concentrated in the thin spherical shell. For MVN distributions with standardized marginal dimensions, this is on the sphere of radius \sqrt{d}.
In fact when I searched "Gaussian curse of dimensionality" I got this at the top result: [1] which shows this phenomena perfectly. Or this one with pretty interactive pictures [2]
[1]: https://www.miryusupov.com/blog/posts/thin-shells/index.html
[2]: https://aseemrb.me/blog/high-dimensional-gaussians-on-a-sphe...