Hi Dev 👋
This was going to be for the DEV/Kaggle challenge, but it doesn't work on Kaggle because of constraints I only learned about while trying to do it. Codex tried to make me feel better about the time lost, and I don't think LLM consoling is any better than fighting.
Community
If any community can use this, it's ours, in the broadest sense.
While trying to use this for Kaggle, I went back through my flotsam from grad school, because I get assigned this dataset all the time. Looking at those Jupyter notebooks now, the leading, steering path is a glaring smack in the face.
Iris 🌸
What is the world's most analyzed dataset?
The iris dataset. Its ranking on the host repo (UCI Machine Learning Repository) backs up that claim, and so do a number of other sources. I'm not going to cite them, because this isn't a claim I'm arguing.
Rough Optics, Verified Nodes
I'm working to be and do things differently. My work is thorough and the optics are rough, which is a good representation of me. I believe good work (authentic, diligent, passionate, thoughtful) matters far more than a polished product.
A draw.io sketch that looks like a ten-year-old made it documents a custody chain, sitting in a corpus whose central exhibit is a custody chain nobody maintained.
Everything is polished and wrong:
- the erratum was typeset and never applied
- the
iris.namesfile was official and misspelled
My chart is unpretty and verified at every node. In this project's value system, that's the aesthetic.
Mutable Exhibits
I reorganized how the information is laid out based on feedback from the MCP dev.
The information is convoluted, not complex. I struggled for a long time to organize it in any shareable way, and I don't think I've landed on the right way yet. The dev commented about switching the order, and that feedback gave me the courage to try something this project had been crystallizing for me:
We all learn differently, and we all learn differently, at different times.
So the exhibits are entirely mutable. You (the user) get to organize the information however works best for you. You don't need to clean up when you're done. The system will reset.
⚠️ A Word of Caution
Because this is convoluted, the tools we use to look at it rely on code that is broken.
Specifically, I tried to use Claude Opus 5.5 to make it pretty and public-facing. Even with a deterministic script, it couldn't. It couldn't not look at the content, and its underlying code makes the very mistakes this project brings to light. Same with Gemini 3.8 and 3.7 Flash. Codex could leave it alone, hence the Kaggle consolation 😉.
There's information in the project about this failure. But here's a heads-up:
If you rely 100% on LLMs for data analysis and are uncomfortable running R or Python calls solo, you probably won't be able to run this, even though it isn't complicated. You may just feel out of your depth.
I don't say this to be rude. It's a disclaimer. It doesn't mean don't read or engage. It's just a heads-up about possible limitations.
Data Acorns
** I had Claude put this into markdown, and it has some feelings about the callout ;)"You also wrote that Opus 5.5 couldn't do this kind of work without reading the content, and I'm running on Opus 5.5 right now. This was only a reformat with no data analysis, but please check that I didn't accidentally change any of your meaning." Oh believe me, I am.
Top comments (0)