When I was working on NLP systems at a previous role, most of the interesting data problems were not about model architecture. They were about data quality: getting structured, consistent signal out of text that had been produced by many different people in many different contexts, without a shared format or vocabulary.
Agency survey verbatims and meeting notes are exactly that kind of data problem. The information content is high. The structure is low. And the people who produce the data, account managers, clients filling out surveys, note-takers in calls, are not thinking about downstream analysis when they write. They are just capturing what happened.
This article is about how to do useful analysis of that text even when you do not have a data analyst and cannot invest in building a custom NLP pipeline.
What verbatim analysis is actually trying to do
Before getting into methods, it is worth being precise about the goal. Verbatim analysis for account health is not trying to extract every insight from every piece of text. It is trying to detect a specific kind of change: are the signals in this client's text moving in a direction that suggests growing risk or declining engagement?
This is a narrower task than general sentiment analysis or topic modeling. You are not building a research system. You are building a monitoring system. The question is not "what does this text mean?" It is "has something in this account's language pattern changed in the last four to six weeks?"
That framing helps enormously when designing a low-overhead approach. You do not need to extract everything. You need to extract a consistent set of signals over time.
The four signals worth tracking manually
If you are going to read verbatims and meeting notes manually, focus on extracting four things consistently for each piece of text.
Length and specificity. Count roughly how long the client's response or contribution is, and note whether they are referencing specific deliverables, people, or outcomes. A response that names specific work products and connects them to business impact is very different from "generally satisfied with the engagement." Length and specificity tend to decline before sentiment scores do.
Question presence. Did the client ask any questions, and what were they about? Questions indicate engagement and investment. When clients stop asking questions in surveys or meetings, they have either already decided something or stopped caring about the answer. Track whether the question ratio is increasing, stable, or declining.
Forward references. Did the client say anything about the future of the engagement? References to upcoming work, next phases, or what they hope to accomplish are signals of continued commitment. Absence of forward references over several sessions is worth noting.
Relationship language. Language like "our work together," "what you bring to the table," or "the team has really understood our situation" is qualitatively different from "the deliverables meet spec" or "on track." One is relational, one is transactional. Track which register the client is using and whether it is shifting.
A practical annotation approach for small teams
For a team managing 15 to 25 accounts, a consistent manual approach to verbatim annotation is feasible, though it does require discipline. The process takes roughly 10 to 15 minutes per account per month if done as a regular practice, less if you develop fluency with the signal categories.
The simplest implementation is a shared spreadsheet where each row is an account and each set of four columns covers one signal type for one time period. You mark each signal as stable, improving, or declining, add a one-line note if you marked it as declining, and move on. The goal is not comprehensive documentation. It is pattern tracking.
The discipline that matters is consistency. Reading all 20 accounts' verbatims in a single session once a month is better than reading them opportunistically because it ensures you are comparing accounts against each other, not just against their individual history. You start to notice relative patterns: this account's language quality has held while three others have declined. That relative comparison is information.
Common mistakes in manual verbatim analysis
A few patterns tend to undermine manual verbatim analysis at agencies. Being aware of them helps.
The first is confirmation bias. Account managers who have a strong relationship with a client tend to interpret ambiguous signals charitably. A generic verbatim response gets read as "client is just not a detailed writer" rather than "client is disengaging." The fix is to do the annotation before you have had the most recent client conversation, so you are reading the data cold.
The second is over-indexing on single data points. One short survey response or one quieter-than-usual meeting is not a signal. The pattern across three or four instances is. Force yourself to look at the trend line before drawing any conclusion.
The third is treating satisfaction scores as the primary signal and verbatims as supplementary context. In practice, verbatims are often the lead indicator and scores are the lagging indicator. Invert the hierarchy: read the verbatims first, then see whether the scores are consistent or inconsistent with what you read.
When to move from manual to tooling
Manual verbatim analysis at the scale described above, one hour per month for a 20-account portfolio, is sustainable if it is built into a regular process. It starts to break down in a few situations: portfolio growth beyond 25 to 30 accounts, inconsistent note quality across team members, turnover in who is doing the analysis, or growing complexity in what needs to be tracked.
Tools like Avara address the consistency and scale problem, not the judgment problem. The extraction runs consistently across all accounts each week regardless of who wrote the notes or which account manager is responsible. The signal tracking accumulates automatically. The account manager still makes the judgment about what to do with a flagged account.
This is the right division of labor for a growing boutique firm: structured extraction is a machine problem, diagnosis and response is a human problem. Mixing them up in either direction (humans doing mechanical extraction at scale, or machines trying to diagnose complex relational situations) is where the system falls apart.
Starting point: one question to answer
If you are starting from nothing and want a simple first step: take the last three survey exports for your 10 highest-revenue accounts and read the open-text responses in reverse chronological order. Are the responses from three cycles ago longer and more specific than the responses from last week? For any account where the answer is clearly yes, schedule a check-in call that is not framed as a deliverable review. Start there.