Skip to Content

2023.11.20

A Developer Conference in Sweden

What are these teams working on?

Aspebodaklungan

I am, for some reason in a hospital?

Changing the name to ??? Meerkats. [Null] Meerkats. Fun, there’s a second team named Girl Sloths. Classic dev hijinks.

Henrkikka Katarina Mattias Annika

Cortex

Something something data lake for medical data neat

Giant Sloth

???

Data Center

Okay neat!

They’re making a data delivery thingy, web platform, okay okay. What is a data delivery system?

Nikita makes a hosting platform which seems cool, like a DO/Vercel/Fly kinda thing for getting containers

Me!

Also John!

He seems neat and fun.

Practicallities

We’re in a hospital. Don’t walk on the white floors because there are MRI machines running and the vibrations will fuck up the images. The screen in this conference room is a 10cm thick glass back projected with 6 high power lasers, with precise color calibration with a 99% black and 100% white. It’s used for medical imaging, radiology and pathology. Don’t touch it or it will make finger prints and cause problems for real patients. We put Google slides on it. It’s unclear why they let us in here.

Annika is Presenting the Pollen Project

The natural history museum in Sweden is responsible for Pollen forecasting. They have a Polynology program for monitoring and data collection, historical and contemporary. The Museum of Natural History here has a large research and pedagogical function, as well as being Swedens largest museum. Neat!

They collect and clean the data, an collaborate with other labs. They have 20 collection stations, and provide data to scientists. The pollen traps themselves are placed on rooftops and things. The collector itself is basically a big sticky tape thing with a fan. The thing sucks up air with pollen, and it sticks to the tape. The tape gets pulled out, died, and is turned into a microscope slide. Then they count the pollen grains by hand. Woah. They have a database they call pollenDB.

In a separate system called StormGeo they create and publish forecasts of allergenic pollens. These are used to let the public know that they are gonna be sneezy next week.

What is Annikas team doing? They work with polleDB which is a Django backend over MariaDB with a jQuery frontend. Woah! Thats probably … okay? StormGeo is external. Their website is SiteVision which seems horrible.

The codebase itself is got submodules (uh oh). The DB itself lives in version control? Woah.

StormGeo is going away, so they’re trying to replace it and allow them to do the forecasts inside of the PollenDB Django app. StormGeo was middlemanning as data broker and reselling the data, so fuck those people. This means that have to do a DB migration to extend the tabes. Yuck! Then they’re making a public API which is nice.

That API acts as a mirror of the MariaDB, so their pollenDB pushes the data into the API’s persistent storage. The API DB can then be read-optimized.

Database tables and models with schemas, relational fields, and schemas suck! I don’t like ORMs! I don’t like database migrations! Fuck that!

Lots of jQuery hate here! But jQuery is at least fucking vanilla JS that doesn’t need transpilation.

The hands on code for the team is making a new UI frontend for the PollenDB system, to make the Pollen experts lives easier. Im into this idea and goal. They are also extending the DB to fit their needs. And then they have to make some components and modules for Sitevision that can talk the API. This includes new charting components as well. Okay! Well!

Challenges

  • Django is new to the team. Django is okay but can be a Thing.
  • The extant codebase is … Legacy™. Standards are … not there!
  • The DB is deep, with data back ton the 70s. It must be big then, and you don’t want to destroy it.
  • Understanding the workflow of the pollen experts. To me, this is the most important part of the entire project.
  • Sitevision is shit because of course it is. “When you develop in Sitevision, you use modern web technologies such as JavaScript and React.” Haha.
  • They have a deadline! Pollen season!

Goals

Most important goal is to make life easier for the scientists. Im interesting in understanding this.

A bonus is to make the data accessible and open.

Ideally, everything is nicer and easier to use for everyone!

So

This is fun, because I know how this is going. I’ve been in the game long enough that I understand exactly what and where the pain points in this process might be, and have ideas on how to mitigate them. There are two sets of interesting questions happening here – one that is strictly technical and looks at how to handle DB migrations, front end migrations, etc etc. The second aspect is about how people work and why they work. Understanding the language of the system is important, as is understanding the interventions our technology can create. The end goal is clearly to make the public sphere more alive, more efficient, and increase human being quality of life. How does the money flow through this system? What does the money enable? Where does human time and energy get snagged and tangled, and how can we divert money and energy to a team like ours efficiently an productively?

A project like this is a public service, and therefor should be treated as such. This means focusing on accessibility and performance.

Nikita Talks About PyScript

As an aside, Nikita has a markdown-to-slides html app running for their presentation – my style of thing to do!

PyScript is WASM bindings for python in the client! I think Observable uses it? No they use Jupyter, which I guess is the framework and tooling for live client/server interactions. PyScript would be much more efficient then. And in fact Jupyter Lite does this for their notebooks! PyScript.com creates a Glitch-like application for writing Python. I like how its isolated, getting python running on a local system can be such a pain in the ass.

Pyodide is the WASM binary, and PyScript then is the custom element wrapper around that. I imagine it can be used to create a REPL which would be fun. They had older one, and their working on a new fancy one.

Pyodide is a 5ish MB payload, but you can use a 250kb micro python binary. That’s not bad! It’s designed for embedded components to write hardware controls in python. That’s neat!

It also has some neat stuff with attribute handlers that attach DOM events to python functions. That’s neat!

You can also import stuff from the standard library. I wonder how that works … dynamic importing? Tree shaking? That’s neat.

  • I’m kind of curious about what the DOM model apis look like.
  • I’m also curious about the JS api interactions.

There some package management stuff that can happen in the browser with TOML configuration. You can bring in NUMPY and marplot and shit, but you need to get the architectural environment you need all the way down the dependency chain, which means C dependencies can get weird. Annoying but workable I guess. Python dependency management continues to be a pain in the ass. It’s like … make sure you reference the right binary distro, cool and fun.

And of course you can run PyScript on a worker thread and interact with it that way, which is the Smart™ thing to do. Then you can use the worker apis to manage that dom/python interaction layer. That’s pretty cool, because all the neat data science stuff is all in Python. Fun!

WASM is good and cool and lets us do good and cool shit.

It’s early development so the api and bugs and stuff are a thing. The weight is wacky and not great, but microPy helps. Additionally, the established front end frameworks don’t always play nice with workers and custom elements which sucks, but that’s on the frontend community for being shitheads.

You can use flask and pyscript! That’s neat.

The main goal here is to lean on all the python libraries that been built up for data science and machine learning. Lots of this work is starting to get into javascript land, but it’s still a specific language with a specific set of goals.

It looks like R has one of these too which is neat, and I know the Rust does this also. It’s a cool world that we’re approaching. It’s interesting because all this shit is just decades of abstraction on top of C, so it kinda makes sense that we’re getting to this point.

I think the approach to using these things in threads, and lazy-loading the binaries after the DOM is done, is a really good and powerful tool. Observable is a good example of how this can be powerful and good.

Workshoppin’ LLM and Stable Diffusion with Erik.

  • Creativity vs Generation
  • Where do we source data
  • Three lines of critique
    • Practical
    • Moral
    • Political

Erik is interested in Creativity and AI – Jürgen Schmidhuber is an influence here. Some discussion about creativity here, some gestures towards some older perhaps ideas like Koestler’s Act of Creation. Inherently, I don’t think current mathematical models be creative, but they can be generative.

Erik asks;

How can we present a problem to automate its solution?

Representations of data and representations of problem spaces is basically the problem. He analogizes this abstract layer of representation with the concept of the compiler, which intermediates between high-level languages and low-level processor architectures. I mean okay!

Maybe we need to learn from data!

Haha okay! Yeah! Ideally! Now we are thinking about deep learning approaches as decomposing complex functions into smaller subsets. This again, is a hard problem. This is what we know how these neural nets work, is the composition of weighted vectors water falling through a set of relatively simple transforms. Yes, composable, functional math is fun an cool!

Where do we get data from?

This is a very key question. When we look at 10billion param data sets, we need to colonize the global south to generate the human labor.

Now we’re thinking about self-supervised learning. Seems like it’s about intentionally degrading data sets and training the model to produce the non-degraded data. The next question here becomes computer labor. How intensive is the compute process on these motherfuckers? How do we feel about that level of compute?

This is about, what do we do to generate models from data sets, but still questions about “where” the data sets come from, and what’s the use and purpose of the generated outputs. Text to image, image to text, yada yada yada. We’re dealing with statistical probabilities that create the hypothetical image, which is a particular fugitive type of hyperimage. These things are actively separate.

I do think that doing things like creating representations in vector spaces and then querying that space for things like clip-image-search and zero-shot classification.

The “generative” part of this use is I think fundamentally unserious, with an unacceptably high cost associate with it.

Stable Diffusion models are also weird and strange, and surprising that they work at all. Training data for these buddies is text/image pairs like images and captions harvested from Flickr or something. Naturally usage rights are an open question here, with a bug ol’ TBD on the legal status of these models. The current US approach is that the state will not protect any intellectual property rights associated with anything generated with these things.

Statistical Learning; observe the data and generate a model that describes the observed outcome and can generate outcomes that are similar to the observed data. This is … linear regression?

LLM do the same thing with word frequency associations. Math is cool and fun sure!

We can use these to generate language

Suuuuuure I guess! See here; stochastic parrots; Thousand Plateaus, etc etc.

ie: A Thousand Plateaus by D&G, and the chapter “Postulates of Linguistics” is I think really interesting as a tool to think about generative AI. D&G tell us that language has two components that relate but are not mapped to each other, but more like a warp and weft of a fabric. One is the content, the other the expression. Both have their own form and topology, that is discrete formal characteristics. To do language you do both together, creating forms of content and layering it with forms of expression. Through this lens, I feel like generative AI, specifically the LLM but also image stuff, is exclusively exploring the topology and form of expression. D&G tell us that language has two components that relate but are not mapped to each other, but more like a warp and weft of a fabric. One is the content, the other the expression. Both have their own form and topology, that is discrete formal characteristics. To do language you do both together, creating forms of content and layering it with forms of expression. Through this lens, I feel like generative AI, specifically the LLM but also image stuff, is exclusively exploring the topology and form of expression. Its got a good handle on the formal qualities of expression, and can manipulate those forms very well. But it’s completely divorced from the forms of content. I think that lets us see how and where it can be useful – for filling in code samples and writing queries, etc etc, we bring our own precise topology of content to the machine and let it explore the form of expression till we get the precise expression we need to elicit the behaviors we want from our other machines

Simon Willison and Andy Baio both have a lot of writing and exploration around this topic that might be worth wrangling, also Ním Daghlian. Gebru and Bender et all obviously need to brought in here and referenced. I am curious what Andy and Simon think about the explicitly ethical dimensions of this, which we’ve discussed a little in person. Its complicate and Baio is quite ambivalent.

There are two basic lines of argumentation here;

  1. What is the practical analysis of the outcome of these tools? Do they produce useful outputs for reasonable inputs?
  2. What is the moral and ethical implication of these models? How do we ethically reason about their creation and use? What is the ethical valence of their outputs?

Both of these lines are argument can be thought of as the act of manifesting values. The practical line questions the practical use-value of the system, the ethical line questions the moral-value of the system.

A third, more tangental or higher-order line of critique lies with the political structures and projects of the people and organizations responsible for these things. This is more of a case by case thing, and Timnit Gebru talks a great deal about the the TESCREAL aspects of OpenAI and Anthropic. TLDR is the political ideologies of these organizations leans nihilist/facist which is obviously a problem.

ChatGPT cannot know what is factual

This is true, they only exist and process in statistical likelihoods! They produce hypothetical outcomes. Sparkling Markov chains.

For WebText dataset for ChatGPT, Reddit was used as a link aggregator which is funny because Reddit has an incredibly particular culture around external links. This that’s to get into the root of practical value argument.

We get around this by having a database of facts. Why not … just query that database of facts? With like … a knowledge graph query engine and reasoner?

Demo?

Stable Diffusion XL from the wackjobs at Stability AI. Something something stable diffusion ui

TL:DR; https://twitter.com/jjellisart/status/1725605342623510967

Practical use

Perhaps as a way to explore and navigate large and complex parametric possibility spaces. “Response Surface Modeling” captures incredibly complex domains.

Erik Ylipää went to art school huh, and is still interested in this shit? Usage as a sketch, usage as exploring hypotheticals.

Using these tools at the beginning of a generative process is tricky, they can be used as foils to help sharpen your own internal embodied knowledge, but they also can short-circuit imagination. I think that there are other mathematical tools that we could use. I would be curious to see how that can happen and be explored.

“Value Delivered” haha yuck.

Summary

For me, this is all about the Image-Cult Society. The dangers here when it comes to “creativity” is summed up cleanly by examine the implications of Graeber&Wengrow and Christopher Alexander. Alexander specifically identifies not-separateness and unique, situated adaptation as essential parts of any loving system. Graeber&Wengrow observe and suggest that the root of societal oppression, slow violence, and the dehumanizing tendencies of the state start with fungibility, the erasure of specific context and community connection in place of the interchangeable commodity.

This is fundamental entirely what stable diffusion and large language models do. They shift language and image into the real of separate, the commodity, and the fungible. There is no specificity or connection, there is no situated context. There is only a flatness of expression, devoid of meaning.

Further Reading

Vanos Talks About The GDI Project

The Genomic Data Infrastructure is a controlled-access database of genomic and phenotypic data across Europe, intended to create a broad base of genome sequencing data.

Lots of shit! Lots of acronyms. Our sibling team is focusing on storage, interface, and something else!

Metaphor; trailer needs a passport and a visa to travel? There’s an auth provider called LS Login which is a SSO auth thing. Then REMS is the system that handles access control and permissions etc etc. Reasonable architecture! No reason to keep those things coupled I guess. REMS provides a JWT that then can be used as a bearer token to the service.

REMS is Resource Entitlement Management System classic. An interesting thing here is management of requests for access.

Beacon is the API that has /info and /query endpoints. What kind of endpoints are those? SQL? SPARQL? The Boolean and count queries with “federated data discovery” sounds like SPARQL to me! Actually no, the boolean query just acknowledges the existence of data sets. There’s a network of beacon apis that federate as services. There are a few different access levels, public, authenticated, and authorized. Neat.

HTSGet is for secure encrypted data streaming of genomic data.

Elsewhere

agile

© 2026