Can and Should a Caption Ever Truly Be Objective?

Written by Payton James, in collaboration with the 2026 Mansfield Training School Memorial and Museum Research Team. Questions or comments may be directed to paytonjames231@gmail.com.

A historical photograph is often understood through the few words placed beneath it. A caption may seem like a simple description of what the photograph shows, but writing one involves more than recording what the camera captured. Every caption requires decisions about how to interpret the photograph, which details to include, what is emphasized, and what must be left out. Each decision shapes how viewers understand both the image and the person depicted in it. So, can a caption ever truly be objective?

Can Description Ever Be Truly Objective?

Consider two ways of captioning the same photograph:

A resident lies on his side on the floor near a brick wall. His knees are bent toward his chest, and his arms are tucked between his legs. He wears a protective helmet, a white shirt, pants, socks, and shoes. A second pair of shoes rests beneath his head.

A resident lies on his side on the floor near a brick wall. His knees are bent toward his chest, and his arms are tucked between his legs. He wears a protective helmet, a white shirt, pants, socks, and shoes. A second pair of shoes rests beneath his head.

A resident lies on his side on the floor near a brick wall. His knees are bent toward his chest, and his arms are tucked between his legs. He wears a protective helmet, a white shirt, pants, socks, and shoes. A second pair of shoes rests beneath his head.

A resident appears isolated as he lies curled tightly on the cold tile floor near a brick wall. He wears a protective helmet and is fully dressed in a white shirt, pants, socks, and shoes. His arms are drawn in tight against his body, knees pulled toward his chest. Beneath his head rests a pair of shoes, used as a makeshift pillow.

Both sentences describe the same photograph of the same person in the same position. However, they leave the reader in two different places. The first is clinical and exact whereas the second, in a way, tells you how to feel about what you are seeing. Neither caption is incorrect, but the two captions serve different purposes.

Every Caption Is a Series of Choices

Those different purposes exist because each caption reflects a different set of choices. A caption cannot include everything. A writer must decide what deserves attention and what does not, and in which order this information will appear. Each choice shapes the meaning of the caption. The moment a writer selects certain details, the caption begins to shape meaning. Blind scholar Georgina Kleege makes a similar argument in More than Meets the Eye: What Blindness Brings to Art while discussing audio descriptions in museums. Although museum audio descriptions often aim for objectivity, Kleege asserts “Even deciding where to begin a description—with the most significant element in the composition or the peripheral details that supply context—requires subjective interpretation” (Kleege 110). The challenge in captioning is not simply choosing the right words, but deciding which details deserve to be included.

The Weight of Words

One of the clearest examples of how word choice shapes meaning is the words themselves. Word choice alone can shift the entire tone of a caption. Consider how differently a sentence can be interpreted based on one word describing the person in the photograph: a resident, a patient, a man, a boy, a child, a victim, an individual. The same is true for verbs such as; look, stare, gaze, watch, glare, observe, or peer. Each verb gives the reader a slightly different understanding of what is taking place. Someone can sit, slump, rest, lean, slouch, sag, perch, or lounge. Some verbs encourage readers to interpret the image in a particular manner.

What a Photograph Usually Cannot Tell Us

Word choices carry so much weight because a photograph itself is incomplete. A photograph rarely captures the full context surrounding an image, including the thoughts, feelings, and motivations of the people within it. The image may not reveal the relationship between people standing beside one another or what happened five minutes, or even five seconds, before the photographer pressed the shutter. These missing details create gaps in the historical record, and some of those gaps remain permanent. Moreover, human nature encourages us to fill in gaps with assumptions and whatever context we do have.

The Risk Runs Both Ways

Filling those gaps, whether through interpretation or through silence, is not without consequence. When writers lean too far into interpretation, they risk misrepresentation, oversimplification, and bias. When writers rely too heavily on flat description, they risk stripping the image of the context that makes it worth preserving. Neither extreme provides a completely safe approach; both interpretation and strict description carry risks. 

Captioning for Accessibility

This same tension between interpretation and omission appears just as sharply in accessibility work. Captions and alt text are not the same tool, even though they are often used interchangeably. The Australian Government’s Style Manual explains, “Alternative text explains information in images for screen reader users. Captions describe images to help users relate them to surrounding text” (“Alt Text, Captions and Titles for Images”). While they serve different purposes, both require interpretation. Alt-text exists to make an image accessible to someone who cannot see it, which means someone still must decide what information matters most to whomever they might imagine will be making use of the alt-text. 

Georgina Kleege explores this tension in More than Meets the Eye: What Blindness Brings to Art, stating that museum audio descriptions often strive for objectivity while inevitably relying on interpretation. Kleege notes, even accessibility guidelines that emphasize neutrality still require decisions about “which details are pertinent” and “how much information is enough” (Kleege 110). Description therefore requires interpretation. 

Interpretive choices must be made, whether for accessibility purposes or not. And whatever gets included, or left out, ends up shaping public memory as much as any other caption may.

Beyond Objectivity

If interpretation cannot be avoided, then the goal is not perfect objectivity, and that may not even be possible. Instead, the challenge is to make thoughtful, well-reasoned choices and to recognize that they are still, ultimately, choices. As Georgina Kleege concludes after examining museum audio description, institutions should “… abandon the pretext of objectivity. It is impossible and beside the point” (Kleege 121). Rather than presenting descriptions as neutral, she argues for acknowledgment of the interpretive decisions they inevitably contain. Every caption reflects the context in which it was written and the person who wrote it, along with who is being imagined as the audience receiving the image. Rather than pretending otherwise, we should be transparent about those limitations.

Works Cited

Kleege, Georgina. “What They Talk About When They Talk About Art.” More than Meets the Eye: What Blindness Brings to Art, Oxford University Press, 2017, pp. 109–121. Oxford Scholarship Online, https://doi.org/10.1093/oso/9780190604356.003.0009.

“Alt Text, Captions and Titles for Images.” Style Manual, Australian Government, 12 Dec. 2024, https://www.stylemanual.gov.au/content-types/images/alt-text-captions-and-titles-images. Accessed 7 Aug. 2026.