Skip to main content

Consulting CV and Project Reference Database: Structure, Tags, and Governance

The data model behind a credentials library that AI can safely retrieve from: CV fields, reference fields, permission levels, tagging, review cycles and GDPR duties.

Florian PloszczykPublished 26 August 202613 min read

Evidence base

Sources behind this article

This article supports its claims with 3 sources. Key sources include:

All 3 sources and access dates

A consulting credentials database has two record types, CVs and project references, and each needs three layers of fields: content, tagging and governance. The governance layer is the one firms skip, and it is the one that decides whether the library is an asset or a liability.

Here is the practical test for whether yours works. Ask it: which of these three project references may I show to a prospect in the German market this quarter, and who confirmed that?

If the answer requires calling two people who were on the project, you have a folder. Let me lay out the data model that turns it into something a proposal team and an AI system can both draw from safely.

Why the structure matters more than the content

Most firms have the content. Somewhere there are CVs, somewhere there are case descriptions. The problem is never that the material does not exist. The problem is that nothing about it is queryable, and that the facts governing its use live in people's heads.

Two consequences follow, and both are expensive.

First, search fails. A consultant spends forty minutes looking for a reference that exists, does not find it, and writes a new one from memory. That new one is now the fourth version of the same project in circulation, each slightly different.

Second, and worse: when you point an AI system at unstructured credentials, it cannot tell an approved claim from an old draft, or a public reference from one cleared only for internal use. It retrieves whatever matches the words. Then it fills the gaps with something plausible, because that is what a generative system does when the retrieval comes back thin.

Structure is what makes retrieval trustworthy. It is unglamorous data work and it is the whole game.

The CV record

Content fields

  • Full name, current role and grade, office and legal entity.
  • Approved biography, in each language you pitch in.
  • Experience statements: the specific claims this person is permitted to make.
  • Education, certifications and professional qualifications, each with a verification status.
  • Languages, with proficiency level.
  • Sector and capability experience.
  • Photograph, with the layout variants your templates use.

The experience statements field deserves attention. It is the difference between a CV database and a free text bio. A statement like "led the post merger integration workstream for a European utility" is a claim the firm will stand behind. Storing it as a discrete, approved item means it can be selected, cited and traced. Storing it inside a paragraph means it gets paraphrased into something slightly different every time it is reused.

Tagging fields

  • Capability tags, from a controlled vocabulary rather than free text.
  • Industry tags, same rule.
  • Project relationships: which project records this person may be associated with.
  • Seniority and role types they can credibly fill in a proposal.

Controlled vocabulary is not bureaucratic fussiness. Free text tags produce "PMI", "post merger integration", "Post Merger" and "integration" as four separate concepts, and your retrieval quality dies quietly.

Governance fields

  • Owner: a named person, usually the practice lead or the individual's counsellor.
  • Approval status: draft, approved, restricted, retired.
  • Last review date and next review date.
  • Availability status, and who maintains it.
  • Consent and geography constraints.
  • Record identifier.

The GDPR position on CV records

I want to be direct about this because it gets waved through more often than it should.

A CV database is personal data about identifiable employees. The GDPR applies in full. That means you need a lawful basis, and consent is frequently the wrong one in an employment context because of the power imbalance. Most firms will rely on another basis, but the point is that somebody has to decide and document it rather than assume.

Beyond the basis, four obligations bite hardest here:

Accuracy. Article 5 requires personal data to be accurate and kept up to date. A CV that still shows a role someone left two years ago is not a housekeeping issue, it is an accuracy problem, and it is also professionally damaging to the person.

Purpose limitation. Credentials collected for proposals should not quietly become an input to performance evaluation or internal ranking. If the purpose changes, reassess it.

Retention. What happens to a CV record when someone leaves the firm? Most firms have never answered this, and the default answer, that it stays forever in the proposal library, is not defensible.

Rights. People can request access, correction and, in some circumstances, erasure. You need to be able to find their data across the library, the proposals it went into, and any retrieval index.

Photographs, availability data and anything performance adjacent deserve extra care. Involve privacy early rather than at audit.

The project reference record

Content fields

  • Client display name, or the cleared anonymised label such as "a European retail bank".
  • Industry and capability.
  • The problem, the approach, the outcome. Three separate fields, not one paragraph.
  • Approved metrics: the specific numbers you may quote.
  • Approved wording: any phrasing the client cleared verbatim.
  • Team members, linked to CV records.
  • Geography and languages.
  • Start date, completion date, and duration.

Separating problem, approach and outcome matters because proposals need them independently. Sometimes you are proving you understand a problem. Sometimes you are proving you can run a method. A single blob of prose serves neither well.

Permission fields

This is the section that prevents the worst kind of incident, and it is the section most libraries lack entirely.

Permission levelWhat it allows
Internal onlyUse inside the firm, never in client facing material
Named under NDAClient may be named in a specific proposal covered by confidentiality
Named in proposalsClient may be named in competitive proposals generally
PublicClient may be named on the website, in marketing and in public speaking

Record who granted the permission, when, in what form, and any expiry or restriction. "The partner said it was fine in 2023" is not a permission record.

Also record the negative case explicitly. A client who has declined permission needs a record saying so, otherwise someone will ask again, or worse, assume.

Governance fields

  • Evidence owner: who can confirm the facts are true.
  • Approval status and approval date.
  • Relevance window: how long this reference stays current for pitching.
  • Review date.
  • Record identifier.

Tagging that actually helps retrieval

Four principles, learned the hard way.

Controlled vocabulary, always. Maintain the tag list as a governed artifact with an owner. Adding a tag should be a small deliberate act, not a free text field.

Tag for the question you will ask. Teams tag by what a project was. Proposals search by what a client needs. Those are different vocabularies. Tag for the second one.

Limit depth. A three level taxonomy that people use beats a seven level one that they abandon. Depth feels rigorous and produces empty branches.

Tag relationships, not just attributes. Which people worked on which projects. Which methodology a project used. Which references support which capability claims. The relationships are where the value is when you need three comparable projects with two available senior people.

Review cycles and expiry

Every record carries a review date. When it passes, the record is unusable until someone confirms it, not merely old.

That distinction is the entire point. A "slightly out of date" reference will get used, because the alternative is more work. An expired one that the system will not retrieve will get updated.

Sensible cadences:

  • CVs: every six months, and immediately on role change, promotion, new certification or departure.
  • Project references: annually, and whenever the client relationship changes.
  • Permissions: at the review date, and whenever a client contact changes, because permission often travels with a relationship rather than with the institution.
  • Metrics: whenever the underlying figure could have moved, and never quoted without a date.

Assign the review work to named people with capacity for it. A quarterly review that lands on whoever is between projects will not happen, because nobody is ever between projects.

How AI should use this library

The library exists so that generation becomes assembly. Four rules make that safe.

Retrieve, never substitute. When no approved record matches the request, the system reports a gap. It does not compose a similar sounding reference from fragments of three others. This is the single most important rule in the entire article.

Carry identifiers into the output. Every claim in a draft deck should trace back to a record. When a partner asks where a number came from, the answer takes five seconds.

Respect permission fields at retrieval time. A reference cleared for internal use only must not surface in a competitive proposal draft, and that has to be enforced by the system rather than caught in review.

Never generate the fields that carry consequence. Client names, outcome metrics, availability and qualifications are retrieved facts, not inferences. A prohibition list in the generation instruction is not optional.

Building it without a two year project

Checklist
  • Start with the CVs of the people who actually get staffed, not the whole firm.
  • Start with the project references that already have documented client permission.
  • Define the controlled vocabulary before entering data, and keep it small.
  • Add governance fields to every record from the first day: owner, status, permission, review date, identifier.
  • Assign owners by name, choosing people who know the truth rather than people with capacity.
  • Load a narrow set, connect retrieval and test it against real proposal requests.
  • Measure three things: right record found, nothing found, and wrong record surfaced to the wrong user.
  • Fix the maintenance loop before expanding: who updates what, when, and who checks.
  • Then expand practice by practice.

That third measurement is the security one. A wrong record surfacing to the wrong user is not a quality issue, it is a confidentiality incident waiting for a bigger audience.

Where offgen fits

CVs and case references in offgen implement this model. Records carry owner, approval status, permission level, review date, market and language. Retrieval runs inside existing engagement permissions, identifiers survive into the generated deck, and an unmatched request produces a marked gap rather than an invention.

More on how this fits proposal and credentials workflows is on our consulting page.

If you take one thing from this: the value is not in having the credentials. Every firm has them. The value is in being able to answer, in seconds and with confidence, which one you may use, for whom, today.

Frequently asked questions

What fields does a consulting CV database need?

Person, role, office, availability, languages, capability tags, the approved biography, the experience statements that person may claim, education and certifications with verification status, photo and layout variant, plus governance fields: owner, approval status, last review date, and consent or geography constraints.

What fields does a project reference need?

Approved client name or cleared anonymised label, industry, capability, problem, approach, outcome, team members, geography, dates, evidence owner, permission level, approved metrics, approved wording, completion date, relevance window and review date. Permission level is the field firms forget and regret.

Is a consulting CV database subject to the GDPR?

Yes. CV records are personal data about identifiable employees. You need a lawful basis, purpose limitation, accuracy, retention limits and a way to handle rights requests. Photographs, availability and performance related fields deserve particular attention, and consent is not always the right basis for an employment context.

How do you stop AI inventing project references?

Retrieval from the controlled library only, a stable record identifier on every claim, a hard prohibition on composing client names or outcomes, and a rule that no approved match produces an explicit gap rather than a plausible substitute. Then the engagement owner verifies anything the client will read.

How should client permission for references be tracked?

As a structured field with levels, not as a memory. Typical levels: internal use only, named use with the client under NDA, named use in competitive proposals, and public use. Record who granted it, when, in what form, and any expiry or restriction attached to it.

How often should credentials be reviewed?

CVs at least every six months and on every role change, promotion or certification. Project references annually, plus whenever a client relationship changes. Set a review date on every record and treat an expired record as unusable rather than as slightly old.

Sources

  1. 01Regulation (EU) 2016/679 (General Data Protection Regulation) EUR-Lex, 2016-04-27. Accessed 26 August 2026.
  2. 02CVs and case references offgen. Accessed 26 August 2026.
  3. 03Consulting industry solutions offgen. Accessed 26 August 2026.

Related articles

Florian Ploszczyk

About the author

Florian Ploszczyk

Co-Founder and COO, MD

Florian writes about consulting workflows, professional presentations, company knowledge, and the controlled adoption of agentic AI.