Visual Elicitation Conversation Analysis
Also known as: VECA, image-elicited conversation analysis, photo elicitation CA, visual-aided CA
Visual elicitation conversation analysis (VECA) is a qualitative hybrid method that uses photographs, drawings, maps, or other visual stimuli to prompt and structure participant talk, and then subjects the resulting interaction to systematic conversation analysis (CA). The approach leverages the evocative power of images to generate richer, more embodied accounts of experience while applying CA's rigorous sequential analysis of turn-taking, repair, and action formation to reveal how meaning is collaboratively constructed in situ.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use visual elicitation conversation analysis when your research question concerns how people interactionally make meaning around a phenomenon that has a visible, material, or experiential dimension — for instance, workplace practices, health decision-making, identity talk, or community memory. The method is especially strong when abstract interview questions produce thin or socially desirable responses, and when the sequential organisation of talk — not just its content — is analytically relevant. Do not use it when the research goal is to measure attitudes or generate population-level frequencies; it is an interpretive, not a quantitative, approach. It is also unsuitable when visual stimuli would be inappropriate or distressing for the participant group without extensive ethical safeguarding.
Strengths & limitations
- Visual stimuli bypass rehearsed narratives and elicit spontaneous, embodied, contextually rich talk that standard interviewing rarely achieves.
- Applying CA's sequential analysis reveals the fine-grained interactional work through which participants collaboratively construct meaning around the images.
- Participant-generated image variants (photo-voice, photo-diary) give participants meaningful agency in shaping what is investigated.
- Well-suited to multimodal settings where gesture, gaze, and artefact handling are integral to the interaction being studied.
- Flexible across disciplines — used in health communication, education, workplace studies, and social gerontology.
- Jefferson transcription and CA sequential analysis are technically demanding and time-consuming; researchers need specialist training.
- Findings are context-specific and are not intended to be statistically generalisable to wider populations.
- The choice and framing of visual stimuli inevitably shapes the talk that is elicited, introducing a methodological layer that must be explicitly reflected on.
- Recording naturally occurring talk around images in situ raises complex research ethics and consent issues, particularly in workplace and clinical settings.
Frequently asked
Do I need training in conversation analysis to use this method?
Yes. CA has a specific analytic vocabulary (adjacency pairs, turn-taking organisation, preference, repair) and requires Jefferson transcription, which takes practice to apply correctly. Researchers new to CA should work through a foundational text such as Sidnell and Stivers (2013) and, where possible, join a data session group before conducting independent analysis.
Who should choose the images — the researcher or the participants?
Both are legitimate and serve different purposes. Researcher-selected images keep the analytic focus on predetermined topics and facilitate comparison across participants. Participant-generated images (photo-voice, photo-diary) foreground participant agency and are better suited to participatory or emancipatory research goals. The decision should be driven by the research question and the ethical relationship you want to establish with participants.
How is this different from photo elicitation alone?
Standard photo elicitation uses images to stimulate interview talk and typically analyses the content or themes of what participants say. VECA adds a CA layer: instead of (or alongside) thematic coding, you analyse the sequential organisation of the talk itself — turn-taking, repair sequences, how participants display understanding or disagreement — treating the interaction as a social accomplishment in its own right, not merely a data source.
What recording equipment do I need?
At minimum, a good-quality audio recorder capable of capturing overlapping speech clearly. For multimodal analysis — where gesture, gaze, and participants' handling of the images matter — a video camera that frames all participants is necessary. In group settings, a fixed wide-angle camera supplemented by a lapel or boundary microphone typically gives the best combination of interactional visibility and audio clarity.
How large a dataset do I need?
CA studies work with collections of instances of a phenomenon rather than with sample size in the survey sense. A robust collection of 20–50 instances of the interactional pattern of interest — drawn from as few as 5–10 recorded sessions — is typically sufficient to support an analytic claim, provided the collection is shown to be systematic and deviant cases are addressed.
Sources
- Harper, D. (2002). Talking about pictures: A case for photo elicitation. Visual Studies, 17(1), 13–26. DOI: 10.1080/14725860220137345 ↗
- Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696–735. DOI: 10.2307/412243 ↗
How to cite this page
ScholarGate. (2026, June 3). Visual Elicitation Conversation Analysis. ScholarGate. https://scholargate.app/en/qualitative/visual-elicitation-conversation-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Conversation AnalysisQualitative↔ compare
- Multimodal Discourse AnalysisLinguistics↔ compare
- Narrative AnalysisQualitative↔ compare