In 2026, the National Gallery of Art in Washington extended automatically generated visual descriptions to more than 140,000 works in its digital collection. Scale immediately changes the accessibility problem: a task requiring years of manual writing can reach almost the entire online museum through one coordinated operation.
For a blind or low-vision person, an undescribed image can resemble a closed door. Title, artist, and medium identify the object, but rarely support a mental representation of it. Description can communicate arrangement, colour, density, gesture, spatial relation, and rhythm, turning a screen reader’s voice into access to visual experience.
The National Gallery states when a description was generated automatically. It has also announced progressive staff review beginning with the most-viewed images and invites users to report errors. These choices recognise automated text as a first editorial version rather than a completed truth.
The page for Mavis Pusey’s Untitled (Abstract) demonstrates the complexity of the operation. Its description identifies black geometric forms and white space, then proposes a city seen from an unusual angle or a collection of puzzle pieces. It records visible elements while constructing metaphor, depth, and familiarity. The machine does more than name; it directs imagination.
The experiment therefore raises a question larger than technical accuracy. Whoever describes an artwork establishes what gets encountered first, which details become important, and where visibility ends as interpretation begins. When this task is automated at museum scale, accessibility also becomes a politics of knowledge.
One Hundred and Forty Thousand Doors
The clearest advantage is coverage. The National Gallery collection contains nearly 160,000 objects, and only a fraction could have received extensive human descriptions quickly. Applying the system to more than 140,000 primary images reduces a historical absence: most digital works acquire a verbal presence.
Scale matters because selective accessibility produces a selective museum. When descriptions cover famous works alone, a blind visitor receives an art history even more concentrated around the canon. Broad coverage makes prints, drawings, photographs, works on paper, and rarely introduced artists available for independent exploration.
Automation also changes the timing of access. Description ceases to arrive as a late addition to an existing page and can become a default part of infrastructure. This shift has cultural value: visitors with visual disabilities become part of the initial audience instead of exceptions served later.
Quantity alone, however, cannot measure accessibility. Text may exist while remaining confusing, inaccurate, generic, or excessively interpretive. Counting described pages measures availability of the service; other indicators are needed to measure the quality of the encounter.
The project should consequently be understood as a threshold change. One hundred and forty thousand automatic descriptions improve upon one hundred and forty thousand silences, provided their provisional status stays visible. The next task is turning automatic coverage into verifiable and continuously correctable knowledge.
Description Already Means Selection
The National Gallery publishes guidelines for descriptions written by people. They recommend beginning with an overview, organising a path through the work’s space, using accessible terms, describing colour, and acknowledging ambiguity when a detail resists identification. Authors are also asked to focus on the web image and avoid knowledge or interpretation imported from elsewhere.
These instructions show that description always involves decisions. In a crowded painting, which element opens the text? Should movement proceed from left to right, foreground to background, or centre to edge? Verbal sequence is linear while an image offers several simultaneous trajectories.
Vocabulary alters perception as well. A line may be rigid, broken, nervous, or dynamic; a surface may seem empty, open, rarefied, or unfinished. Each word distributes emotion and intention. Useful objectivity separates observation from inference instead of pretending that language is transparent.
In abstract art, this tension becomes especially visible. A system must make relations among unnamed forms memorable. It often uses resemblance to architecture, landscape, bodies, or familiar objects. Metaphor can support a mental map, while also fixing a figure that the work had left unstable.
Description therefore builds an order of access. The central AI problem arises less from interpretation itself, since people also interpret, than from difficulty locating where interpretation began and which data directed it. An automatic label reveals origin but still leaves observation, analogy, and hypothesis merged.
Mavis Pusey and Automatic Metaphor
On the page for Untitled (Abstract), the automated description begins with a verifiable foundation: black geometric shapes, angular edges, straight lines, and white spaces. These details provide orientation, contrast, and structure. A person unable to see the image can begin assembling a mental composition.
The text then moves toward resemblance. Forms evoke a skyline from an unusual viewpoint or irregular puzzle pieces. Neither image simply resides inside the work. They are interpretive bridges translating abstraction into recognisable objects and influencing what the listener attempts to imagine.
The description also introduces overlap, depth, and dimension. Overlap may be a graphic relationship, while depth is a perceptual effect. A system able to separate these levels could first state how forms touch, cross, or remain apart, then explain that their arrangement may suggest spatial depth.
The answer is never complete removal of metaphor. Many visual-description practices recognise that a small interpretive component can make artwork vivid and intelligible. Difficulty begins when an analogy carries the same certainty as a colour or count, making hypothesis sound like an object’s property.
Pusey’s page therefore becomes an editorial laboratory. It could offer three selectable layers: factual summary, extended description, and possible spatial or metaphorical readings. Users would receive greater access because they would also know the status of the words forming their image.
The Asymmetry of Verification
A sighted person can read the description and compare it immediately with the image. If the system invents an object, misreads a posture, or assigns depth to a flat surface, contrast remains available. For a blind person, the text may occupy the position of principal witness.
This difference creates an epistemic asymmetry. Automated error does more than convey incorrect information; it can become the artwork as experienced. A visitor might discuss a composition, remember it, and connect it to other images through an element that never existed, while lacking an independent channel for verification.
Transparency must therefore extend beyond the sentence “automatically generated.” It should indicate whether the text awaits review, has been checked by staff, or has been evaluated by a blind or low-vision consultant. Review date, revision history, and correction status become part of accessibility.
Feedback must also work through assistive technology. A generic button may be technically reachable yet require too many steps or prevent precise reporting. An effective process could classify omission, invention, ambiguous wording, identity error, and excessive interpretation.
Blind and low-vision users should never function merely as recipients or unpaid error detectors. System design, evaluation, and governance require their continuous and compensated participation. Accessibility reaches maturity when people who depend on descriptions also help define their quality.
When Error Becomes Metadata
An online description quickly exceeds its screen-reader function. Search engines index it; educators and students copy it; translation tools transform it; other systems may summarise or absorb it. Text written as an accessibility layer can become a general source about the artwork’s content.
This circulation amplifies error. A shape identified as a building can influence subject searches; uncertain gender can enter a record; a threatening atmosphere can orient later readings. Algorithmic interpretation gains stability through repetition.
Metadata determines museum visibility. Titles, subjects, keywords, and descriptions influence which works emerge from search and which stay hidden. If automated text feeds recommendation systems, a single reading can become an ordering criterion across the collection.
Accessible description should consequently remain distinct from verified curatorial data. Its text may be searchable, but every reuse should preserve provenance and review status. Copying its sentences while dropping the label converts an identifiable draft into an anonymous assertion.
Version history would turn error into useful evidence. Showing that a description changed, together with the reason, could build trust without claims of infallibility. The institution becomes accountable by making recognition and repair observable.
Inequality of Review
According to the project announcement, staff plan to review all descriptions while beginning with the most-viewed images. The choice is understandable: scrutiny starts where an error would reach the largest audience. It is an efficient strategy for reducing aggregate exposure quickly.
It also produces a temporary hierarchy. Masterpieces and already popular pages receive human verification first; unfamiliar works, marginalised artists, and rarely searched materials remain automated for longer. Popularity begins to determine the initial quality of access.
The mechanism may reinforce the canon. Frequently visited works receive reliable descriptions, making them easier to use in classes and research, which generates more visits and additional priority. Less visible collection areas risk retaining generic text precisely where greater contextual care would have the most value.
A balanced policy would combine several criteria: traffic, likelihood of error, visual complexity, significance to particular communities, representation of lesser-known artists, and user reports. Random sampling would also measure performance beyond popular pages.
Success should be reported through legible indicators: percentage reviewed, average correction time, error categories, distribution across media and collections, and evaluations by blind and low-vision users. The project’s scale would gain a second metric grounded in care rather than coverage alone.
The Museum as Accountable Narrator
The National Gallery experiment cannot be reduced to replacing curators with a machine. It shows a museum using automation to address an otherwise enormous accessibility backlog. Origin disclosure, planned human review, and a feedback channel are meaningful forms of responsibility.
The next step is layered description. Users could choose a brief version, detailed visual route, curatorial context, or interpretive reading. They could also discover language, date, model, review status, and human contributors without turning an artwork page into a technical report.
The museum should preserve plurality. For complex works, one description risks becoming the image’s official verbal equivalent. Alternatives written by artists, educators, and blind or low-vision contributors could demonstrate that different perspectives can coexist with factual accuracy.
Contestability becomes an accessibility feature. Knowing a text’s origin, reporting a problem, following a correction, and choosing the degree of interpretation provides access to the editorial authority constructing the verbal artwork. Visitors gain tools for participation as well as a description.
Automating 140,000 descriptions opens the museum; governing them will determine which museum opens. Artificial intelligence can transform a silent collection into a narrated landscape, but every narration establishes distances, hierarchies, and metaphors. Institutional responsibility lies in making the machine a declared, correctable, and plural first voice, so access to art never requires uncritical trust in a single synthetic gaze.