LLM Research (6 blogmarks)

← Blogmarks

Attention is all you have

https://alicegg.tech/2026/09/21/attention

I can tell that I've had my head in LLM research papers a lot lately because I saw the title of this post ("Attention is all you have") and thought it looked a lot like Attention Is All You Need.

And that is related to opening illustration of the blog post which says that if you focus on something enough in a short span of time, you start to see that thing everywhere and in everything.

Anyway, what the post is really about is reclaiming the aspects of the internet where we are in control of our attention instead of the ones that hijack it.

The good thing is that this intentional internet is still around. It has just been a bit buried below the corporate web, but it’s not very hard to find. After all you’re on this blog, so you probably already have a good idea about it.

Moravec's paradox

https://en.wikipedia.org/wiki/Moravec%27s_paradox

Moravec's paradox is the observation that, as Hans Moravec wrote in 1988, "it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility".

I hear the description of this paradox referenced fairly often, but I didn't know it was a named paradox attributed to someone (Hans Moravec).

The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study

https://arxiv.org/html/2605.23135v1

Supervisory Engineering Work

We propose a new category of work we term supervisory engineering work, encompassing the direction, evaluation, and correction of AI output.

The productivity-experience paradox of working with AI-coding assistants:

We also identified a productivity-experience paradox: productivity perceptions held stable, with 84% reporting improvement at both time points, yet among matched participants, the proportion reporting worsened developer experience in at least one dimension nearly doubled from 14% to 27%, with flow state and cognitive load eroding while feedback loops improved. These findings suggest that AI coding assistants are impacting both the nature of software engineering work and how engineers experience it.

There have been studies on other aspects of LLM using in software as well as productivity for students, but not of software professionals.

Further, what remains largely missing is an understanding of how professional software engineers experience these tools, and whether those experiences change over time.

It strikes me that we need a classification system for the different forms of AI-assisted software development. These tools are used in widely non-uniform ways. From tab-completion for the current line or next couple lines to series of small prompts where the developer views and confirms each changeset to “auto mode” flows where an entire task or feature is completed by an agent harness to multi-loop agentic flows and beyond. Productivity gains and developer experience must vary widely across these types of AI development.

LLM tooling can tighten the feedback loop during development.

However, the usefulness of this feedback depends on its quality and relevance, and developers still report needing to validate AI output carefully.

Cognitive Fragmentation:

Overall, prior work indicates that AI coding assistants can improve aspects of developer experience while simultaneously introducing new sources of cognitive effort and fragmentation.

Concerns about maintainability had a big jump between the initial survey and subsequent survey. Which I interpret as the more these tools are used, the more maintenance becomes a growing concern.

Quality was the dominant primary concern at both time points, though it declined from 44% to 36%, followed by security. The most notable shift was in maintainability, which rose from 3% to 19% as a primary concern between Q1 and Q2.

Learning, onboarding, and code understanding are improved by LLMs and coding harnesses.

Sub-theme 2B: Accelerated learning, onboarding, and code understanding. Participants in both time points used AI to accelerate understanding across unfamiliar languages, frameworks, or legacy codebases. For example, one participant wrote: “Much easier to pick up new tools, libraries, languages and get productive.” (P132, Q1). Another wrote: “It has been particularly helpful in explaining legacy code and detecting bugs more efficiently.” (P2, Q2). Others reported gaining conceptual clarity: “Definitely felt like it’s given me a level up on understanding new concepts” (P5, Q2).

This supervisory work doesn’t really map onto existing tasks. It’s often compared, in my experience, to being a lead who is reviewing and giving feedback to others on the team, but I don’t think this is quite the same activity.

This supervisory work replaced portions of hands-on implementation effort but does not map cleanly onto traditional software development task categories. As one participant explained, “Sometimes, I’ll not be happy with the original reference, and will ’coach’ the AI to produce something that works better for my use-case”

Interesting, there is a “Use of Gen AI” disclaimer at the end of the paper.

We used Claude Code (Anthropic) to assist with developing R scripts for statistical analysis and reviewing written text for clarity and readability. All outputs were critically reviewed and validated, and responsibility for the correctness of the analysis and interpretations rests entirely with the researchers.

Trying to wrap my head around watermarking LLM-generated text

Laurie Voss wrote No, Claude's watermark doesn't make the writing worse which is probably the most accessible starting point. If visuals help, Laurie pointed to this interactive watermarking post.

The idea and research behind the watermarking solution being proposed by Anthropic is SynthID Text which huggingface introduces here with code examples for getting started with it.

This post from Anthropic does a good job of explaining the whole thing in a nutshell in this paragraph. The core gist being that while probabilities or odds of what is generated are unchanged, this doesn't impact the model's weights, but rather "randomness decider" uses an internally-known key rather than an RNG.

Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude.

Inventing Transformers

https://alexcbecker.net/blog/inventing-transformers.html

This post is a survey of several papers and techniques from 2012 to 2017 that were influential in the development of the modern Transformer Architecture. Which in turn is what is behind the rapid advancement in Large Language Models.

I like surveys like this because they give a high-level overview that can contextualize and point at both things I've already read about like Word2Vec as well as things that are new to me like Adam.

The 2025 AI Engineering Reading List

https://www.latent.space/p/2025-papers

The people at latent.space have curated a list of ~50 papers across ten areas for AI engineers looking to dig into relevant research in the AI and LLM space in 2025.

Here we curate “required reads” for the AI engineer. Our design goals are:

  • pick ~50 papers (~one1 a week for a year), optional extras. Arbitrary constraint.
  • tell you why this paper matters instead of just name drop without helpful context
  • be very practical for the AI Engineer; no time wasted on Attention is All You Need, bc 1) everyone else already starts there, 2) most won’t really need it at work

Funny to see that they completely side-step "Attention is All You Need". As foundational as that paper is, they consider it old news and not practical for engineers in 2025.

There is a repo LLM Practical Guide full of papers and an "LLM Tree" diagram that may have been a partial source of inspiration for the latent.space paper clubs. Unfortunately, this repo hasn't been updated since 2023. That said, there are still a lot of good resources in there.

This "LLM Tree" diagram is excellent showing the different branches of LLM tooling over time.

LLM Tree