repsens.AI

RepSensAI — AI Disclosure & Automated Decision-Making Statement

Effective: [EFFECTIVE DATE] · Version: 1.0 Published at https://repsens.ai/ai-disclosure

This document describes how RepSensAI uses artificial intelligence and statistical scoring, what the outputs mean, and what they cannot be used for. It is written for three audiences: the dealership deploying the platform, its HR and legal advisers, and the salespeople being measured.

It is provided so a deploying employer can meet its own obligations under laws governing automated decision-making, profiling, and automated employment decision tools. Provider is the developer of this system; the dealership is the deployer. Where such a law requires a developer to give a deployer documentation of a system's purpose, inputs, limitations, and known risks, this document is that documentation.


1. What the system does

RepSensAI measures how a salesperson performs and trains them to improve. It produces:

  1. Performance metrics — arithmetic on activity and sales data (e.g. closing rate = sales ÷ ups). Not AI.
  2. An effectiveness index — a weighted composite of the metrics that can be measured. Statistical, not AI.
  3. Skill marks (training, current generation) — a large language model marks each written criterion of a practice exercise as met, not yet, or cannot be determined, quoting the words it relied on. These roll up into mastery bands — Learning, Practicing, Retained — never a score and never a peer ranking.
  4. Drill grades (training, prior generation) — a large language model grades a full practice conversation against a fixed rubric, producing a drill score and a 0–99 rating that combines practice performance with real production. This generation runs alongside the current one while the two are compared.
  5. Coaching text and daily briefs — generated by a large language model.
  6. Manager and owner reports — narrative summaries generated by a large language model.

2. What the system does NOT do

3. What goes in

Performance inputs — units sold, ups, leads assigned, appointments set and shown, demos, write-ups, and activity counts (calls, texts, emails). Sourced from the dealership's CRM or DMS exports, or self-reported by the salesperson where no feed covers that field.

Training inputs — the salesperson's typed or spoken responses during practice drills.

Configuration inputs — the store's own goals and standards, set by the dealership's management.

Not used as inputs — protected characteristics; commission, earnings, or personal budgets; demographic or inferred data of any kind.

4. How the rating is calculated

4.1 Metrics. Each funnel metric is computed arithmetically and compared against a target that the dealership sets. Example defaults: 40 contacts per day; a 20% up-close rate (sales ÷ ups).

4.2 The availability rule — the most important design decision. A metric is shown, scored, and coached on only if the underlying data exists. If a store's systems cannot supply a denominator, the metric reads "not tracked" and is excluded from the index, which is then re-normalized across whatever remains.

There is a second gate: a metric that exists but has too small a sample is displayed honestly (for example "n=2 of 4") but does not become a verdict and is not scored.

Consequence, stated plainly: two salespeople at different stores may be scored on different metrics, because their stores measure different things. Scores are meaningful within a store, and comparisons across stores should be treated with caution.

4.3 Skill marks — the current training generation. A skill states two to four written criteria. After a practice exchange, a language model marks each one met, not yet, or cannot be determined, and must quote the salesperson's own words as evidence for the mark. There is no score: a criterion is either demonstrated or it isn't.

Cannot be determined is a judgement about our material, not the person. The exercise is re-served at most twice; if the mark still can't be made, the content is flagged as defective and the attempt is never counted against the salesperson.

Marks accumulate into a band per skill — Learning, Practicing, Retained. Bands are not numbers, are not averaged, and are never ranked against other salespeople. Reaching Retained requires clean repetitions spread over days, on varied scenarios, including at least one after a gap and one mixed in among other skills — because repeating one exercise until it passes measures recall, not skill. A repeated failure moves a band back down; it does not lock anyone out, and the next clean repetition restores it.

Certification honesty. Typed practice can raise a salesperson no higher than Foundations Proficient. Any statement that someone is ready for live customers requires a manager's validation — a human observation, recorded with the method used (observed roleplay, reviewed call, live shadow, outcome review, or an explicit override). The software cannot award this to itself.

4.3b Drill grading — the prior training generation. Each drill is graded on five dimensions defined by that drill's scenario card, each on a 1–5 scale with written anchor descriptions for 1, 3, and 5. The five grades are summed and multiplied by four, giving a drill score of 20–100. Grading is performed by a large language model that receives the rubric, the scenario, and the transcript. This generation remains available while the two are compared; §4.4 and §4.5 describe it and apply to it alone.

4.4 Subject ratings (prior generation only). Drill scores roll into per-subject ratings using an exponentially weighted moving average, so recent performance counts for more. Ratings decay over time without practice, and increases are capped by the difficulty of the drill attempted, so easy drills cannot inflate a rating.

4.5 The overall rating (prior generation only). Combines subject ratings with the salesperson's real production. A subject never practised is treated as an unproven floor (25) rather than an average — an unproven skill lowers the overall rating rather than inflating it. This is a deliberate conservative choice and it means a low rating may reflect a lack of practice rather than a lack of ability.

4.6 Narrative text. Coaching text, daily briefs, and manager reports are generated by a large language model from the metrics above. The model is instructed to work only from supplied numbers, but generated text may still contain errors and should be verified against the underlying figures.

5. Known limitations and risks

Stated directly, because a deployer needs them for its own assessment.

LimitationWhat it means in practice
Data quality drives everythingA wrong, late, or misconfigured CRM/DMS export produces wrong scores. Nothing downstream can detect this.
Unequal measurement between storesDifferent stores support different metrics, so scores are not strictly comparable across rooftops.
Practice ≠ the floorDrill performance measures a simulated conversation. It is a proxy for skill, not a measurement of real selling.
LLM marking variesThe same transcript may be marked slightly differently on different runs. The current generation reduces this by asking for a yes/no judgement on one written criterion at a time — the shape a language model judges most reliably — and by requiring quoted evidence for every mark; it does not eliminate it. Treat single marks as evidence, not proof.
Bands are not a person's abilityA band describes what the records show about one narrow skill practised in writing. It is not a rating of the salesperson, and no band is a statement that they are ready for live customers — only a recorded manager validation is.
Low rating ≠ low abilityUnpractised subjects sit at the rookie floor. A salesperson who does not use the drills will score low regardless of talent.
Assignment effectsMetrics based on assigned leads reflect how leads were distributed as much as how the salesperson worked them. Lead distribution is controlled by management.
Proxy-discrimination riskAlthough protected characteristics are not inputs, activity-based metrics can correlate with schedule, shift, tenure, language, or lead assignment — which can correlate with protected characteristics. A deployer that uses these outputs in consequential decisions should monitor for disparate impact. We will supply the data needed to do so.
Generated text can be wrongNarrative summaries may misstate or over-interpret figures. The figures govern.

6. Required human oversight

Every dealership using RepSensAI agrees in writing that it will:

Where the law of a deployer's jurisdiction requires an independent bias audit of an automated employment decision tool, that obligation falls on the employer using the tool. We will provide the methodology documentation, input definitions, and output data an auditor needs. Contact legal@repsens.ai.

7. Voice features

Where voice practice is enabled, speech is converted to text by the user's own web browser. We receive and store the text transcript and timing measurements only. We do not receive or store audio, and we do not generate voiceprints or any other biometric identifier.

This is a deliberate architectural choice: by never holding audio, the system does not process biometric identifiers of the kind regulated by the Illinois Biometric Information Privacy Act, the Texas Capture or Use of Biometric Identifier Act, or comparable statutes.

If this ever changes — if audio is stored, or a voice is used to identify a person — this disclosure and the underlying analysis must be revisited before launch.

Timing-derived measurements (for example, pause length after a price is presented, or the ratio of speaking to listening) are measurements of conversational behavior, not biometric identification. They cannot be used to identify anyone.

8. How a salesperson can question a result

  1. Raise it with your manager. They can see the underlying numbers and correct store data.
  2. Ask for the inputs. You are entitled to know which metrics were scored and which read "not tracked."
  3. Request a re-grade. A drill grade can be reviewed. Grading varies slightly between runs, and a result that looks wrong may be.
  4. Contact us at privacy@repsens.ai for questions about the data itself. Questions about how your employer uses the results belong with your employer.

No adverse action should follow from a score alone. If you believe one has, raise it with your employer's HR — and you may tell us at legal@repsens.ai, because a deployer using our outputs that way is in breach of our agreement with them.

9. AI providers

Large language model processing is performed by Anthropic, PBC in the United States. We have contracted for terms under which data submitted for processing is not used to train Anthropic's models. We do not train any model on customer data.

10. Changes

Material changes to scoring methodology will be reflected here with a new version and effective date, and dealership customers will be notified.

Questions: legal@repsens.ai