AI can analyse open-text employee survey feedback by finding recurring ideas, grouping related responses, and drafting concise themes. But AI does not make comments anonymous by itself. Privacy depends on what text enters the process, who can access it, how model output is checked, and whether the final report exposes individual wording or only findings supported by several independent responses.
The safest pattern is therefore not comments in, summary out. It is a controlled pipeline: restricted collection, eligible-group checks, anonymous evidence references, thematic analysis, deterministic validation, minimum support, and employer-facing findings that cannot be traced back to one response.
What does open-text survey analysis actually do?
Open-text questions let employees explain why a score is high or low, describe friction that a fixed scale did not anticipate, and suggest changes in their own terms. The analytical task is to turn that varied language into a small set of useful patterns without mistaking the most vivid sentence for the most common problem.
Traditional qualitative analysis does this through coding and theme development. Researchers become familiar with the material, mark meaningful features, group related codes, review candidate themes, define them, and produce an account of the patterns. Braun and Clarke's widely used framework describes thematic analysis as a flexible method rather than a single mechanical formula. Braun and Clarke's thematic-analysis paper
AI can accelerate parts of that work. A language model can compare differently worded responses, recognise that several comments concern the same underlying issue, suggest a broader label, and draft a concise description. For example, responses about late roster changes, unpredictable overtime, and difficulty planning family commitments may all contribute to a theme about scheduling instability.
That proposal is analysis, not proof. The system still needs to establish which independent responses support the theme and whether it is safe to publish.
Why are raw survey comments hard to anonymise?
A name is only one way to identify a person. A response may mention the only night shift, a distinctive client incident, a rare job title, a recent disagreement, or a phrase colleagues recognise. Even a perfectly ordinary sentence can become identifying when a manager combines it with workplace knowledge.
The UK Information Commissioner's Office explains that removing direct identifiers is not enough when data can still be linked with other information. Its guidance highlights singling out and linkability: whether someone can isolate a person's record or connect it to what is already known about them. ICO guidance on effective anonymisation
NIST reaches the same practical conclusion from a de-identification perspective. Its guidance covers free-format text as well as structured records and warns that some apparently de-identified data can be re-identified. NIST guidance on de-identifying personal information
For employee surveys, that creates four distinct risks:
- Direct identification: a name, email address, phone number, employee number, or link appears in the text.
- Contextual identification: a role, event, location, shift, or project is unique enough to reveal the author.
- Language identification: recognisable wording, spelling, or a copied sentence points to one person.
- Evidence isolation: a theme is based on only one or two responses, so the report effectively exposes an individual's concern even after paraphrasing it.
Deleting names addresses only the first risk.
What does a privacy-safe AI analysis pipeline look like?
A useful design separates source material from publishable output. The model assists with interpretation inside a restricted process; it does not decide what HR is allowed to see.
| Stage | What happens | Privacy control |
|---|---|---|
| 1. Collect | Employees submit scores and written answers | Source responses remain outside employer-facing reports |
| 2. Establish scope | The system determines which organisation or reporting groups have enough responses | Small or unsafe scopes are suppressed before reporting |
| 3. Label evidence | Each eligible written answer receives a temporary, non-identifying reference | The model can cite support without receiving participant identity as its evidence key |
| 4. Propose themes | AI groups similar meaning and drafts broader findings | Instructions prohibit quoting and identifying detail |
| 5. Verify output | Application code checks structure, evidence references, identifiers, and source-language overlap | Invalid or overly similar output is rejected rather than published |
| 6. Enforce support | The application counts distinct evidence references behind each finding | A finding below the reporting minimum is withheld |
| 7. Publish | HR sees the theme, broad sentiment, and support count | Individual comments and participant-linked answers do not cross the report boundary |
This division of responsibility matters. Models are good at proposing semantic groupings, but application code is better suited to enforcing fixed rules. A model should not be trusted to invent its own support count, decide that one dramatic response is important enough to publish, or remember every privacy rule consistently.
How do comments become supported findings?
Consider a survey that asks what makes day-to-day work harder. The source responses may use very different language:
- Several people describe frequent priority changes.
- Others describe unfinished work being displaced by urgent requests.
- A smaller set mention unclear ownership when plans change.
- One person describes a highly specific incident involving a recognisable colleague.
A thematic model might propose Unstable priorities are disrupting delivery as a candidate finding. It should also return the anonymous evidence references that genuinely support that idea.
The application can then make deterministic decisions:
- Are all returned evidence references real and part of the eligible input set?
- Do at least the required number of distinct responses support the finding?
- Does the title or description contain a direct identifier?
- Does it repeat a recognisable sequence of words from any source response?
- Is the result a broad pattern rather than a uniquely identifying incident?
Only after those checks should the finding enter the employer-visible results layer. The specific incident can still inform the broader pattern where appropriate, but its wording and identifying details should not appear in the report.
Why should support counts come from evidence rather than the model?
Generative AI can produce fluent statements that are wrong. NIST calls this risk confabulation: confidently generated content that is erroneous, internally inconsistent, or unsupported by the input. NIST's Generative AI Risk Management Profile
In survey analysis, a confabulated number is especially dangerous because it changes the apparent importance of a workplace issue. A report that says 12 people raised workload pressure must be backed by 12 distinct valid evidence references. The number should be calculated by the application from those references, not copied from prose generated by the model.
The same rule applies to prevalence labels such as “widespread,” “isolated,” or “most employees.” A model may help draft language, but quantitative claims should come from deterministic counts and eligible aggregate data.
What should HR receive instead of individual comments?
A privacy-safe written-feedback section can still be useful. It should answer:
- What recurring issue or strength appeared?
- How many distinct answers supported it?
- Was the broad pattern positive, mixed, or negative?
- Which survey question or protected reporting scope produced it?
- What follow-up question or action would help leadership understand it better?
It should not reveal:
- The original sentence or a lightly edited quotation.
- A distinctive project, incident, shift, role, or personal detail.
- A finding supported by fewer responses than the fixed minimum.
- A timestamp that can be matched to completion activity.
- A filter or export that lets HR reconstruct the underlying response.
This is a deliberate trade-off. Leaders lose the colour of a quotable sentence, but employees keep a much clearer protection: their individual wording does not become management material.
What can AI get wrong when analysing survey text?
Privacy controls do not make the analysis infallible. Common quality risks include:
Combining different problems
Two responses may share vocabulary while describing different causes. “Too many meetings” and “not enough communication” both concern communication, but combining them into one theme may hide a meaningful contradiction.
Splitting one pattern too narrowly
A model may create separate themes for workload, deadlines, capacity, and overtime even when the evidence describes one broader resourcing problem. This can make the report look busier while weakening each support count.
Losing minority nuance
A repeated theme is not automatically the only important theme. Serious safety, harassment, or grievance disclosures may require a separate confidential reporting route rather than being exposed through an anonymous culture report or discarded as an unsupported finding.
Overstating sentiment
Sarcasm, mixed feelings, and conditional statements are difficult to classify reliably. Sentiment should be treated as a broad aid to interpretation, not a psychological measurement of the author.
Reproducing source language
Even when instructed to paraphrase, a model can echo distinctive phrases. Automated overlap checks and rejection are necessary because a prompt is not an access control.
For these reasons, an AI-generated survey finding should be traceable to protected evidence inside the restricted process, testable against fixed rules, and open to appropriate quality review without exposing source comments to employer users.
How does Candora analyse written feedback?
Candora keeps participant operations and employer reporting separate. Identifiable participant records support private invitations, one-response enforcement, completion tracking, and reporting-group attribution. Submitted answers are held in a restricted source layer that employer accounts cannot access.
After a survey closes, Candora first creates privacy-safe aggregate results and eligible reporting scopes. Written answers for an eligible scope are then processed by a restricted findings job. Each answer receives a temporary evidence label such as comment_1; the label is local to the analysis and does not contain the participant's identity.
The AI proposes thematic findings and attaches only those evidence labels that support each one. Candora then verifies that every label is valid, rejects direct identifiers, rejects output that repeats any sequence of three or more words from a source answer, and calculates support from distinct evidence labels. A finding needs support from at least three distinct written answers before it can be published.
HR receives the resulting paraphrased finding, broad sentiment, tags, and mention count. It does not receive the source comments, response identifiers, timestamps, or a way to recut the report until an individual answer is exposed. This is why Candora describes the output as privacy-safe findings, not anonymous comments.
The boundary is also precise about what AI does and does not do. The restricted processor does read the source text to analyse it. The promise is that individual written answers and participant-linked responses remain unavailable to the employer, while only validated, sufficiently supported findings enter employer-facing reports.
What should you ask an AI survey provider?
Use questions that reveal the complete pipeline rather than the quality of a demo summary.
- Can employer users open, search, or export individual written responses?
- Can an administrator turn anonymity off for one survey?
- Does the AI receive participant names, emails, or other identity fields?
- How are direct identifiers and distinctive source wording kept out of generated findings?
- Must every finding cite valid internal evidence references?
- Is the support count calculated by application code or asserted by the model?
- What minimum number of independent responses is required before a finding appears?
- Are small reporting groups excluded before written feedback is analysed for that scope?
- Can HR combine filters or exports to recover the comments behind a finding?
- What happens when model output fails a privacy or evidence check?
- Which AI provider processes the text, where is it processed, and what retention terms apply?
- Is there a separate confidential route for urgent issues that need individual follow-up?
A trustworthy answer should describe data, roles, thresholds, and failure handling. “Our AI anonymises it” is not enough.
Is AI analysis safer than manual comment review?
It can be, but only when the surrounding architecture removes the employer's access to source comments and enforces the output rules. Giving HR an AI summary while retaining a “view original” button leaves the central identification risk in place.
Manual qualitative analysis can also be responsible when trained analysts work in a restricted environment, generalise identifying detail, verify themes against several responses, and publish only protected output. The useful comparison is not human versus AI. It is controlled analysis with a fixed privacy boundary versus unrestricted access to individual comments.
AI makes recurring analysis faster. The architecture determines whether it is safe.
To see the difference between employer-anonymous and confidential survey models, read Confidential vs anonymous employee surveys. For the complete Candora data boundary, see How Candora protects employee survey responses.
quick answers
Frequently asked questions
How does AI analyse open-text employee survey responses?
AI can group responses that express similar ideas, propose concise themes, and classify broad sentiment. A trustworthy system then verifies each theme against distinct source responses, applies a reporting minimum, and publishes a paraphrased finding rather than treating the model output as evidence by itself.
Does removing names make survey comments anonymous?
No. A comment may still identify its author through a project, role, shift, incident, location, or recognisable wording. Privacy controls must address indirect identification and the employer's access to source comments, not just delete name fields.
Should HR receive raw anonymous survey comments?
Raw comments can be useful but carry a high identification risk, especially in small or well-known teams. A safer employer report presents recurring, paraphrased findings supported by several independent responses and withholds isolated observations.
Can an AI model guarantee employee anonymity?
No. An AI model can help classify and summarise text, but anonymity depends on the complete data and reporting architecture around it. Access controls, eligible reporting groups, output validation, support thresholds, and the absence of raw-comment views are what protect the employer-facing boundary.
How does Candora protect written employee feedback?
Candora processes written answers in a restricted source layer after a survey closes. Employer accounts receive only paraphrased findings that pass identifier and source-language checks and are supported by at least three distinct answers; they cannot access individual comments or participant-linked responses.



