Your project is already in your archive
How a small arts organisation used AI to connect scattered content and help visitors discover more.
In the dark days before Large Language Models (LLMs) seeped into the public consciousness, introducing new technologies into your tech stack would almost always have been needs-driven, and application-specific. That is to say, an encountered problem would have triggered a search for a solution, and then, the identified solution would define the integration requirements, you having to adapt your systems to allow for the promised streamlining.
AI, and that’s a broad brush, as a tool, turns this approach somewhat upside down. Introducing ‘AI’ is not a solution in itself. Instead, its general applicability means that, if you’re interested in deploying AI to attempt to streamline a part of your production process, it’s you (or your team) who has to make the decision as to what problem you’d like to focus on, followed by how you design a solution to tackle it. Instead of the parameters of the implementation being defined by the problem and the tool, now you can pick the problem, as well as the solution, and its implementation. AI is not a panacea, but it can be used on many fronts, in many ways.
You will also be able to decide how closely connected to your production pipeline, your Content Management System, you want this new integration to be. You can put the AI outside of your existing system, alongside it, or you can work on tight integration, and they each have their own advantages and risks.
But you can also consider putting this fancy newish tool, AI, completely outside of your news production pipeline.
You’re probably not at an organisation that operates at the scale of the New York Times, but as a benchmark, it’s useful, as they produce around 60000 articles per year. That quickly escalates, making meaningful access to archival content a challenge.
Of course you will have made sure, just as the 700, or so, tech workers at the Times, that your content can be easily searched through. But the challenge is not around when a user already knows what it is they are looking for. That problem was solved a while ago. The challenge is surfacing relevant content when the user doesn’t know they were going to be interested in what you’ve queued up for them.
If your archives are very well organised, that is to say, very well categorised along a number of orthogonal taxonomies, constructing an approximation of a user’s interest based on behaviour, and matching that interest with what’s in your archives, is ‘just’ a matter of doing a few matrix calculations.
But, for many of us, this fine grained categorisation will be absent, or inconsistently applied.
Over the last 25 years, I’ve worked with predominantly small news organisations in the Global South where, often, I was the only tech specialist in the company. Your organisation might be a bit larger, but you likely won’t have the 700 techies at your disposal which the New York Times can rely on.
Being small, and having very limited resources, requires being smart. A kind of applying the 80/20 rule, and making sure that the 20% of functionality that is expected to take 80% of the time is picked such that, in reality, you can work without that 20% of functionality, leaving yourself with 100% of what you need, in just 20% of the time.
One part of ‘being smart’, is the clever use of automation, and in the context of improving categorisation of historical content in your archives, the obvious approach to unlock its hidden wealth. Yet, given conventional tools, some classifications are not too easy to implement; a location-taxonomy will struggle with identical place names, and many place names can also be people’s names. So, just running a matching algorithm with a long list of places will not be sufficient.
One of the things I do, is run an NGO for artists. It’s called walk · listen · create, and it’s the home of walking artists and artist walkers. You can imagine it’s a tad niche.
That said, as a network organisation, we have nearly 9000 registered users, and facilitate over 3000 contributing artists from all over the world, if centred on Europe and North America.
We got there, in part, by pulling in the material that members of the network produce or are involved in, which includes walking pieces, events, videos, books, and more, and then making these accessible in a single, centralised location, through a unified interface, the Museum of Walking.
This has worked well, hence the accumulation of over 3000 contributing members, but also has meant that, by the end of last year, we had a loosely organised archive going on some 15000 items. Admittedly, we’re no New York Times, but my whole organisation is only 4 part time volunteers, including myself, while our content is also extremely unstructured, exactly because it’s produced by 1000s of creators, all with very different ideas on what matters, what needs attention, or, even, how things should be called.
In short, this had created a bit of a mess.
I try to balance the downsides of AI with its opportunities. Thankfully, AI is not NFTs, that is to say, it’s not only hype and get-rich-quick schemes. And, even though we’re being lead to believe that AI is much more capable than it really is, there are many areas in which deploying AI-assisted technologies, with some care, can be hugely beneficial.
That said, though the world seems to have moved on from the issues around AI copyright abuse, the jury is still out on to what extent society will accept the abusive energy needed by the data centres required to satisfy the insatiable hunger of Large Language Models.
Nevertheless, one of the low hanging fruits for which AI is pleasantly capable, is categorising and summarising content. At its core, an LLM, after all, is little more than a stochastic parrot. And now, quite a sophisticated stochastic parrot, meaning that it’s quite good at reducing material that you throw at it to a reasonably credible core of subjects, or topics. Or places, or whatever taxonomy you’d like to measure your content against.
So, late last year, I started implementing processes that would take the material in our archives, one by one, and thoroughly classify each item. On relevance, categories, and location. Though I first reduced the user-generated taxonomy of some 10.000 terms, to a more manageable 3000, and then extracted categories that were, in reality, places.
Around 95% of this work is now automated, handled by carefully designed prompts which have grouped taxonomy terms, identified irrelevant material in the archive, categorised the remainder, and associated them with the most relevant locations.
The result has been a much more tightly integrated archive, where visitors might arrive at any single piece in the archive, but then, through natural progression, can find themselves on a path through material that is related to the content they were visiting. It’s like unlocking the long tail of our content.
This has been a success; since late last year, visits and page views have about tripled, the majority of the increase being exploration of the archive.
And we didn’t need the resources of the New York Times to get there.
A major advantage of putting AI to work outside of your production pipeline, is the relatively low risk involved. You’re not producing content, news, with AI, and so hallucination is not likely to result in reputational damage. A badly categorised piece of content, when presented to a user, will simply be seen as exactly that, if at all. That is to say, badly categorised, no harm done.
This optimisation happens in a very low-risk, high reward environment.
On top of that, if you keep destructive changes to a minimum, everything you do can be reversed, if and when needed, also if you only discover a mistake months down the line, or if you decide a change is necessary in how you want to organise your taxonomies.
The full AI analysis of our archives isn’t yet quite finished. The 5% of manual intervention I mentioned above is the quality control that I’ve built into the process. I could do without, but as I have found that even my carefully designed robots need an occasional tap on the fingers, I’d prefer to keep the overall rollout a bit slower than theoretically possible, if that means I can safeguard myself from mistakes that might be obvious to me, but less so to a parrot, and thus save myself the consequential additional work.
Yes, designing for this manual intervention, a bit of a combination of human-in-the-loop and human-on-the-loop, does mean the process takes a bit more time, also because my time is limited.
But, we have to have priorities. I can spend my time only once, and I don’t always want to sit behind a big screen and deal with improving an archive. Sometimes, I like being able to go on an artsy walk myself, for example by learning about the abusive past of the Dutch colonial project in Brazil, centred on Recife, where I found myself last week.