Tomorrow we will have a discussion in our lab about how to use AI as a scientist. I have increasingly started to use AI, having moved from Copilot Chat that our university provides last year, to some free credit on GitHub Copilot under their educational program, to Claude Pro integrated in Visual Studio Code, to a Claude 20x max subscription (in part to also make these AI tools available to my lab in case they want to use them).
I work with collaborators whose opinions and values are important to me, and several of them are more skeptical about the use of AI. I see their concerns, both in how my students use AI, and how some peers use AI. I am as annoyed by people use write texts where AI has made up references as anyone. I also see how tempting it is to use AI mindlessly. It does so much, so quickly, and validating that it has done things well is so effortful, that you get sucked into overtrusting AI. Some people don’t want to use AI at all, but for me, the question is not whether to use AI, but how to use it well. Of course, I don’t see any problems with how I use AI, and like in the rest of my work, I need others to criticize me to point out my own biases. So I am writing down my current thoughts, to make them as explicit as possible.
Using AI as a scientist
matches my personality and work
I use smart
home tools, and have been using If This Than That (IFTTT) for years. I like to
use technology to make my life more efficient. In the last year, I have been
using AI to create tools that improve my life. For example, I have asked an AI
to create a pipeline where I can transform text (books and papers) to audio on
my laptop, sync it to my phone, where I have a custom made Android app that I
can use to listen to these audio files. When I swipe back, I can mark the
section I listen to, and the text (with the use of meta-data files that link
the time stamps in the audio to the sentences in the text) of this section gets
automatically synced to my Obsidian note-taking tool. If that feels like
overdoing it to you, and that you might go as far as buying a tablet to read,
but that building custom software is a bridge too far, I understand. My
personality is such that I like to be efficient. I would predict that the more
efficiency matters to you in your job, the more you use technology, including
AI.
AI use also aligns with work I do. I spend a lot of time behind my computer, interacting with software, working on websites, and I program R packages and tools such as Metacheck. Of course a computer already automates a ton of our activities, so AI use is an extension of the tasks we automate. But if you spend less time behind a computer, and especially if you do not code, there will be less use-cases for AI.
When and how to use
automation?
I taught a Human Factors course, and the book we used has a chapter on human-automation interaction. It starts with the question “Why automate” and offers 4 reasons: 1) it is impossible or hazardous for people to perform these tasks, 2) tasks are difficult or unpleasant, 3) automation can extend human capability, or 4) we do it just because it is possible, even where it has no benefits. To the extent that we see AI as automation, I have never used it to do impossible or hazardous tasks, but I think I have experienced the other 3 categories. Sometimes the use of tools has been useful, sometimes less so (but learning when AI use is useful is for me just part of learning to use new technology).
Doing tasks that are
difficult or unpleasant
Let’s start with unpleasant tasks. I often perform tasks that take a long time, and/or need to be monitored, but that require very little cognitive resources, and that an AI can do.
Grading
I teach
Introduction to Psychology and Technology in the bachelor this quarter, and
Metascience in the master. Both courses have weekly assignments that should be
graded. The colleague who was supposed to assist with grading fell ill. We use
the Canvas learning management system. Many of the assignments are mainly
intended to activate students – for example, they have to walk around campus
and take pictures of things in the built environment to assist blind and/or
deaf people to find their way, or they perform an online experiment and have to
copy-paste their results. All I do for these assignments is check if they have
completed them, and then I assign them the points. There are also open
questions that need actual grading by comparing answers to a rubric.
Canvas has
an API, but figuring out how it works is too difficult and unpleasant so I
never used it in the past. But AI can figure out how the API works. I use AI to
write R scripts that download all assignments, and store the questions and
answers in a spreadsheet. I use that spreadsheet to grade, which prevents me
from having to click 20 times for all 160 students in the slow online Canvas
system. When I am done with grading, I use an R script to upload the results to
Canvas.
I also use
Claude to check if assignments are completed. In the past, I would give all
students a point for some questions without checking if they actually uploaded
real pictures, as the assignment instructed. I knew some students uploaded
images unrelated to the assignment in an attempt to make it look like they did
the assignment, and I disliked that I did not have the time to grade these
assignments, but I found it an acceptable trade-off for more engaging class
assignments. Now, I can use AI to check the assignments I would otherwise just
assign points for, thereby keeping the more engaging assignments, while also
checking that students who try to get a point without doing the assignment receive
no points.
There are a
number of similar use-cases for grading. I have compared the quality of grading
of low-stakes assignments last year, compared to this year, and the use of AI
has improved the accuracy of the grades I assign. To be clear, I was aware that
my grading was too lenient, but I think weekly graded assignments are more
educational than removing them due to a lack of time to grade them carefully.
For both courses, the assignments that are graded with AI are a small
percentage of the final grade, unlikely to determine who passes a course, and
who not. The psychology course has an in person final exam and for the
Metascience course I have switched to (rather labor intensive) oral exams to
make that master course AI proof.
The time saved on grading with the use of AI is substantial (approximately 2 hours per week, for 8 weeks), and the quality has increased. At the same time, I invest some of this time back into education, making new and more engaging assignments, and more personal assignments and tests. I think it has made my education better.
Monitoring
I perform many tasks on a computer that take a long time. In the past, I ran simulations that would last days, and currently, I run automated checks that Metacheck performs on thousands of scientific articles. When a computer runs for days, there are many things that can go wrong. AI helps me to create scripts that are more reliable, that store intermediate results and can restart easily, informs me when there are unforeseen problems, and allows me to quickly resolve them. Claude allows you to chat with another computer through the browser, which makes it very easy to leave a computer at work on over the weekend, while monitoring its progress, and asking Claude to resolve problems that emerge. For example, if during a run it becomes clear that there is a bug in a script, Claude can notice this before the run has continued for days, fix the code, and restart. I remember how in the past I would often rerun the same thing for multiple times across multiple days (often a real source of frustration in my work).
Adopting new Best
Practices
I try to
improve my research practices, but there is only so much I am willing to learn.
I have been using GitHub, as an amateur, but at least I managed to made code
open and have version control. But I never received training in Git, and I have
used it suboptimally. Claude makes it much easier to adopt good coding
practices, as well as cleaning up a repository after I made a mess of things.
Creating new branches, using Large File Storage, creating public releases,
making code reproducible, and documentation have improved a lot.
I am also
increasingly using Docker, which was another tool I knew I had to learn to use,
but found difficult. I was able to run containers in the past if there was a
clear tutorial, but now I can also make docker containers and share them (for
example as part of Metacheck’s reproducibility_check where we automatically
check if code in repositories is computationally reproducible in a docker
environment where most R or Python packages are pre-installed to speed up the
checks).
In an ongoing project, where I am iteratively improving the code I am using, and therefore the data I am generating is updated, I was losing track of which files where updated, and which not, with the risk that the manuscript described stale results. I knew targets was the tool I should use to prevent this confusion, but I really did not have the time to learn it. With the help of Claude, I reorganized my repo and implemented it anyway.
Extending Human
Capabilities
There are
only so many things I can learn, and my lack of knowledge limits the things I
can do. This is especially true when it comes to writing software and code.
AI helps me
to complete ideas I have that I would otherwise never get around to complete.
For example, I turned static pictures (generated by code I had written and validated that performs the correct calculations underlying these code-generated pictures in earlier versions of the textbook) into interactive visualizations. I want free
open textbooks to be better than paid textbooks by commercial publishers, and
interactive apps are one way to achieve this. But I don’t have the time to learn
how to turn my code into an interactive app. AI is very good at this. This is
also true for the creation of Shiny apps. It took me a long time to learn Shiny
at a mediocre level, but with AI, turning code into an R Shiny app is trivially
easy. I have to load Shiny apps on my server,
which runs on Linux, but I don’t know Linux well, and AI is also of great help
here. Shiny apps are now more up to date, because maintaining them and updating
them takes less time.
One area where the use of AI becomes more problematic is when ‘extending human capabilities’ means ‘doing more in less time’. Where scientists used to have to bike to the university computer, hand in their programming cards, and return a day later to get the results, we can now perform statistics on our laptop in seconds. Where we used to browse through paper journals, we can now download any pdf version of an article we want. These innovations mean we can do things more quickly. Where should we use AI to speed up our work, even if we can also do this work manually? And, how do we know we are extending our capabilities, instead of using AI to create flawed work?
Validating code written
with AI
When I let
AI write code, I extensively validate the code against ground truth. This is
easy when I know what results code should return, and not possible when I do
not know what code should return. When I program on Metacheck and I develop a
new module, for example one that automatically retrieves datafiles from
repositories, I use manually coded ground truth (so, a datafile where people coded the correct information as accurately as possible) to validate the code I create
with AI. I know that my code should find repositories, and I can check if it
actually finds them. Even in this validating step, AI can be very helpful.
Where I might manually check 20 repositories, I can use AI to check all
repositories.
Here, we
get to some challenges I experience when using AI. Manual validation against
ground truth sounds like a good golden standard. But it is effortful, so using
AI to validate AI is tempting. It is less work, and in all fairness, we are
also fallible human beings, so sometimes our manually coded ground truth is
wrong, and AI points this out. I think I spent more time validating code as I
did before I used AI, but validation efforts are more extensive and more
systematic, and they often occur at a higher level. I can feel a greater
distance between the code and data and my validation work, towards an
evaluation of the output of an entire piece of code.
Still,
where needed, I can easily jump back in, and look at the dataframes my code
generates, to see if the contents match my expectations. This requires coding
skills, so AI can’t replace learning to code (even if it might change how we
learn to code). For most of what I am building, AI seems to be doing very well,
so this is increasingly less necessary – or maybe I am overtrusting my AI, and
I think it is not necessary.
I spent a
huge amount of time validating code (as I spend less time coding, I spend more
time validating, and for many of the modules in Metacheck that I have created I
spend weeks (often hundreds of hours) validating that they work, and
iteratively improving them. For this reason, I don’t like the term
‘vibe-coding’. I feel this is a pejorative term that is used when people throw
together code in a short amount of time without knowing what their code does,
and whether it works. This is not what I do when I code with AI.
When I
started coding, I did everything manually. Then I moved to copying code from
stack overflow, then to copy-pasting code from ChatGPT, and now I ask Claude to
create the code. This is an improvement over asking ChatGPT (and if you have
not used AI inside your programming software, you don’t really known how coding
with AI works and can’t judge it). Because Claude has access to the code, and
the results, it can check the input, code, and output. If you have a good
default prompt (e.g., you tell it to always check everything, never make
assumptions, etc.) the quality of the code is hundreds of times better (and
really not comparable) to asking ChatGPT ‘how do I create a dataframe with data
organized like I want’.
I also
think there is a lot to learn about how to set up a workflow where I am even
more likely to notice when AI created code that is not in line with what I
intend to create. I am just learning how to use these tools – of course in a
decade I will be much better at it. When I learned to share my data openly, I
linked to a dropbox folder in my scientific articles. Of course these links are
dead now, and this was really bad practice. I am sure I am making similar
mistakes when I use AI. But the mistakes might already be smaller than when I
figured out new technology (like data sharing) myself.
When there
is no ground truth, I would be much more careful in using AI. I also use AI to
write reproducible manuscripts where I do not know what the results should be. In
these cases, I don’t know if a number should be high, or low. I might still ask
AI to create a first version of the calculation of the descriptive report. But
then I have to double check every number, and here, I often intended to compute
other numbers than AI thinks I wanted to generate. It is still helpful, and the
corrections are less work than writing all reproducible analysis code from
scratch, but the required oversight is much larger, and I really need to look
at the raw data and code to make sure everything is accurate. Here, I will
still use AI to ask how to compute certain output, then I check the code, I
correct it where needed, but then I would also ask AI to check my corrections.
After all, I also make mistakes, and AI can help to catch these more quickly.
So, there is a lot of contextual variation in how much I use AI, and what role it plays in the work I am doing. I don’t spend less time checking my work – if anything, I spend more time on it – but the nature of the checks have moved to a slightly higher level, often more away from the data and individual code lines, to the results that the code should produce. This is actually not that different from when I used to code simulations manually. I would also evaluate the results on a higher level (the results of the simulation) and only jump into the dataframes and individual lines if I expected that there was something wrong. At the same time, those simulations were built in smaller steps, and the individual sections of code would receive at least a basic check. Now, these lower level checks are still present, but instead of manual checks, I often ask AI to write explicit tests (e.g., in this dataframe, column 1 should be a number between 0 and 1). This is on the one hand a better practice, but sometimes also comes closer to letting AI check AI. It’s a space where I still need to explore how to create a reliable and efficient workflow.
Using AI for Writing
As
enthusiastic I am about using AI for code, as disappointing I am in using AI to
write. Maybe it’s because I have a lot of experience in writing, but I find AI
written text mostly mediocre. This also holds for Peer Review. The only use I find
in AI for peer review is (beyond maybe using Claude to automatically reproduce
your results) is to remove the Dutch Directness that can creep in some reviews.
There are
some exceptions. Sometimes AI can improve parts of text I write. I write too
long sentences, and sometimes as a non-native English speaker I create weirdly
formatted sentence constructions, I don’t know the right words, or the thoughts
just appear on paper somewhat unclear. Asking AI to suggest an alternative
sometimes (but not always) provides inspiration for how to improve the
sentence.
Another
exception is low-stakes writing. Sometimes as academics we need to create text
that will not be read (or not be read by many people), where we would otherwise
invest so little effort that our writing would not be better than what AI
produces. Of course, you could also say ‘why not reduce bureaucracy in
institutions’? I agree. With many things, AI use amplifies underlying problems.
If people use AI to write, maybe what they are writing does not need to exist.
If AI increases our reflection on this topic, that is a good thing.
Another exception is documentation. If I want to summarize the code in a repository, AI will do so more diligently that I can. Similarly, the Metacheck Manual has large sections that are written by AI (although I also rewrote a lot of sections) because it summarizes what code does, and AI can do this relatively well. Ideally I would write every line myself, but I don’t have the time, and the entire team is busy making software. So, AI makes documenting things less work, and although it is certainly not as good as what a human would produce, given our priorities and limited resources, it suffices. If you prefer manually written documentation, as I do, and you have time, it would be a great way to help with Metacheck.
Other uses of AI
I personally like the ability to automate boring tasks (e.g., asking AI to prepare a meeting, copy files from my computer to a harddrive, cleaning up messy folders, look at which of 10 meetings dates is free in my calendar, etc.). I also like the voice interface of Claude. Instead of reading Wikipedia (which I often do), I can ask Claude questions while am walking around. The answers largely come from public sources like Wikipedia, but I can more quickly switch and develop my thoughts about a topic I don’t know a lot about. I don’t like the use of AI to generate pictures, videos, or presentations. I find them bland and uninteresting. But, this drifts away from the specific use of AI as a scientist, and this blog is long enough as it is, so I will not go into details about other uses of AI too much.
Ethical aspects of AI
I am regularly
reflecting on the ethical aspects of using AI. I am not an expert on this, so I
can’t say sensible things. I hope there will soon be a European Union version
of Claude for scientists, and I will switch over to one if it exists. I prefer to
spend my money on non-commercial and open access tools. I am hopeful this will
soon become possible. But in the meantime, I think Anthropic is not that
different from my bank. There are a lot of values I disagree with, but the
benefits of having a bank outweigh the problems I have with banks in general. I
am supportive of the general policy in the EU around the use and development of
AI, so I hope that in the future I will be happier with the organizations that
provide me access to AI than I am with banks. I do not have a drivers license,
but if you do, you should probably be more upset about the companies behind
the oil you buy than about the company that provides your AI inference, in my
humble opinion.
Then there is the climate. Data servers use a lot of energy. Here, the question is how much energy we want to use, and what we want to use it on. I think people using AI to generate or adjust pictures and videos should stop using AI. I don’t see any societal benefit, and it uses quite some energy. I also think we should reduce our reliance on data servers where possible. I personally have stopped automatically syncing everything to the cloud. For example, I maintain a local back-up of my pictures, but I no longer store them in a datacenter. If we want to reduce the use of datacenters, I would personally prioritize a reduced use of video streaming. I feel services like Instagram reels waste everyone's time and are overly addictive, and I would not mind it if people who want to use them need to pay the environmental cost (compared to the current free use) of streaming, and I think the world is better off if these services do not exist. I also think we should stop streaming video with services such as Netflix. As someone who likes vinyl, I think we should own the media we consume. If we want to rely less on datacenters, I would prefer a societal movement away from all streaming services such as Netflix. As video streaming takes up a much larger percentage of the use of datacenters, I think for my own personal life, I prefer to use energy more on AI, and less on things like video streaming.
This blog
is mainly an attempt to make my current behavior and beliefs more explicit, so
that they become easier to criticize. I am sure I am getting things wrong, I am
missing things I should weigh more, I have not discussed all important aspects,
and in 10 years I will read this and know better. But I at least feel prepared
to go into our lab meeting next Monday, and spend some time with the
old-fashioned use of a keyboard to organize my thoughts. (Yes, this 100% human
writing but it would be funny if you thought otherwise).