Methodology
Anonymity is not a promise. It's a property of the system.
This is the complete methodology behind Company Radar — how it's built, how results are generated, where its limits are, and exactly what it does and does not claim to know. Every section below expands into a full explanation, a concrete example from the report itself, and the specific limitations that apply.
01Anonymity Architecture
Identity and response are never linked in the database — not encrypted, not pseudonymized, structurally absent. There is no field, index, or join path connecting a participant to their answers, at any point after submission. This is a property of the schema, not a policy applied on top of it.
+
Anonymity Architecture
Identity and response are never linked in the database — not encrypted, not pseudonymized, structurally absent. There is no field, index, or join path connecting a participant to their answers, at any point after submission. This is a property of the schema, not a policy applied on top of it.
Detailed Explanation
Most tools that claim anonymity implement it as a promise enforced by process: an administrator agrees not to look, or a field is encrypted with a key someone, somewhere, could still access. Klarwerk's approach is different in kind. When a participant is invited, a row is created holding their email, a single-use access token, and their department. When that participant submits their assessment, a separate set of rows is created — one per answer — holding only the panel ID, the question ID, the dimension ID, and the score. Critically, this second set of rows contains no field that references the participant record at all. Not a foreign key, not an anonymized hash, not a token. The column simply does not exist in the schema.
This matters because it changes the nature of the guarantee. A system that encrypts the link between identity and response is still, in principle, reversible — by the organization operating it, by a court order, by a future engineer who didn't know better. A system where the link was never created cannot be reversed by anyone, because there is nothing to decrypt, subpoena, or accidentally expose. The distinction is not academic: it's the difference between 'we choose not to look' and 'there is nothing to see.'
The token a participant receives is single-use and tied only to their invitation record, not to any response. Once they submit, their invitation record is marked complete — but that completion flag lives entirely within the invitation table, updated independently of whatever they answered. An administrator can see that participant X has completed the assessment. They cannot see, derive, or infer what participant X said, because the system was never built with a path from one fact to the other.
Aggregation compounds this protection. Even if the identity-response link existed (it doesn't), a single response would still be reported only as part of a group. Scores are calculated by averaging across all responses to a given question within a given segment — company-wide, or department-level once a department clears the minimum group size. No individual response is ever surfaced, displayed, or exported on its own, at any stage of the pipeline, by any role, including Klarwerk's own administrators.
This architecture was a deliberate response to the central failure mode of internal HR-run surveys: employees correctly suspect that an internal tool, run by people who report to leadership, cannot fully guarantee that answers won't eventually be traced back to them — and they moderate their honesty accordingly. Removing that suspicion requires more than a stated policy. It requires making the thing people are afraid of technically impossible, so that trusting the process doesn't depend on trusting the people running it.
In the Dashboard & Report
On the survey itself, every screen displays a compact 'anonymity architecture' strip restating the four load-bearing facts: identity and answers are stored separately, no manager sees individual responses, results are reported only above a five-response threshold, and free text is never shown verbatim. This isn't decorative — it's the same claim made on this page, repeated at the exact moment a participant is deciding how honestly to answer.
Limitations & Confidence
Structural anonymity at the individual level does not eliminate every re-identification risk. In a very small department, a participant may still be able to infer something about the group's aggregate answer even without seeing anyone's individual response — and in departments near the minimum threshold, elimination reasoning ("I know what I said, and the aggregate suggests...") becomes marginally easier as group size shrinks. Klarwerk mitigates this with the five-response minimum and recommends against creating departments smaller than eight people, but this is a mathematical property of small-group statistics, not a flaw specific to this system, and no aggregation threshold can fully eliminate it for a two- or three-person team.
At a Glance
Why Executives Should Care
A workforce that doesn't trust anonymity answers every future assessment with the same caution it uses in a performance review — and every signal this platform produces degrades with it. If this guarantee is ever perceived as broken, response honesty drops immediately and doesn't recover on its own; the next cycle's data becomes unusable for exactly the decisions it exists to inform. A credible anonymity architecture is what makes it possible to detect a retention or alignment risk while it is still cheap to fix, rather than after it has already cost you a team.
Executive Case Study
- Situation
- A 60-person logistics company launches its first Company Radar cycle. Three employees privately ask their manager whether responses can really not be traced back to them.
- Observation
- The manager forwards the question rather than answering it personally. Response rate finishes at 85%, and Trust & Psychological Safety scores 20 points below the company's other dimensions.
- Interpretation
- The low Trust & Safety score is plausible and specific, not an artifact of participants being afraid of the tool itself — the architecture held, and the finding reflects the organization, not doubt about the survey.
- Business Implication
- Leadership can act on the Trust & Safety finding with confidence, rather than spending a cycle wondering whether the number reflects reality or fear of the instrument.
- Recommended Executive Action
- Cite the technical anonymity architecture explicitly in the next internal communication, rather than repeating a general reassurance — specificity is what converts skepticism into participation.
Business Impact
- Cost if this signal deteriorates
- Every future assessment cycle silently loses honesty, converting a diagnostic instrument into an expensive confirmation of what leadership already believed.
- Executive risks that increase
- Leadership continues operating on a materially incomplete picture of organizational risk, with no way to know the picture is incomplete.
- Decisions that become harder
- Any resourcing decision built on a report becomes harder to defend, since the underlying data quality can no longer be assumed.
- Typically affected
- Response honesty across every subsequent cycle, not just the current one.
Where You Will See This
02Survey Design
Klarwerk's 40 questions are deliberately built around one design constraint: every item must force recall of a specific, rememberable instance rather than invite a general feeling or trend judgment. This is what separates the instrument from a standard engagement survey, where abstract agree/disagree statements are the norm.
+
Survey Design
Klarwerk's 40 questions are deliberately built around one design constraint: every item must force recall of a specific, rememberable instance rather than invite a general feeling or trend judgment. This is what separates the instrument from a standard engagement survey, where abstract agree/disagree statements are the norm.
Detailed Explanation
The instrument consists of 32 fixed scale questions (four per dimension, across eight dimensions) and 8 open-ended questions (one per dimension), answered on a 1–5 agreement scale. Every question went through the same test before being included: could this be answered from a vague impression, or does it require the respondent to check it against an actual memory? Early drafts of several questions failed this test — for example, an early version of a Leadership item asked whether 'confidence in leadership's judgment increases rather than decreases' during difficult periods, which is a directional trend judgment answerable without recalling anything specific. The final version instead asks respondents to think of 'the last genuinely difficult period this organization went through' and assess whether leadership's judgment held up — a concrete anchor that produces a materially different, more honest answer.
This discipline was applied uniformly across all eight dimensions, not just the ones that received early design attention. Strategy & Alignment's item on cross-team priorities, for instance, doesn't ask whether teams 'pursue the same priorities' in the abstract — it asks whether two teams have, at some point, worked toward the same priority in ways that actually pulled against each other. That's a concrete, falsifiable claim a respondent either has or hasn't witnessed, not a general sentiment they can answer on autopilot.
Ten of the 32 scale questions are reverse-scored: agreement with the statement indicates an organizational risk, not organizational health. Processes & Collaboration's items are entirely reverse-scored — for example, agreeing that 'when an official process fails, people create workarounds instead of fixing the process itself' is a negative signal, not a positive one. This is intentional: framing every question in the same positive direction would make the instrument easier to skim and answer without engagement, and would also make it trivially easy for a respondent to fall into a response-set pattern (agreeing with everything, or disagreeing with everything) without the scoring engine detecting it. Mixing polarity forces genuine engagement with each item.
The eight open-ended questions exist for a different purpose than the scale items: they capture the specific texture of an issue that a 1–5 score cannot. Where a scale item might establish that psychological safety is low, the corresponding open question — 'describe, in general terms, the kind of issue someone here would think twice before raising' — surfaces what kind of issue, in the respondent's own words, without ever naming a person, team, or event. These responses are never shown verbatim to anyone; they exist solely to identify recurring themes once aggregated across five or more responses.
The full instrument is fixed across every customer and every assessment cycle. No question is customized, reordered, or removed per company. This is a deliberate trade-off: a fixed instrument sacrifices some flexibility in exchange for something more valuable at scale — comparability. A department's Trust & Psychological Safety score means the same thing this cycle as it did last cycle, and eventually, the same thing across the customer base as a whole, precisely because nothing about the question changed in between.
In the Dashboard & Report
Each question screen in the live survey shows, directly beneath the question, a 'Scientific rationale' note and a 'Why are we asking this?' note — the same design discipline described here, made visible to the person answering, not just documented after the fact. A participant answering the Trust & Safety item about junior employees challenging senior colleagues sees, in real time, that this targets hierarchical candor asymmetry as a leading indicator of organizational risk.
Limitations & Confidence
A fixed, non-adaptive instrument cannot capture every organization's specific context. A company with an unusual structure — no formal departments, a fully remote workforce, a recent merger — will still answer the same 40 questions as every other company, and some items may land as less directly applicable to their situation than others. Klarwerk accepts this trade-off deliberately in exchange for cross-cycle and, eventually, cross-company comparability; a fully adaptive instrument would produce richer individual context at the cost of any ability to compare results over time or against a benchmark.
At a Glance
Why Executives Should Care
A vaguely-worded instrument produces vaguely-actionable data — scores that move a few points between cycles without telling you whether anything real changed. Concrete, memory-anchored questions are what convert a survey result into a specific executive decision instead of a mood reading. This directly affects decision quality and execution speed: the sharper the input, the less time leadership spends debating whether a finding is real before it can act on it.
Executive Case Study
- Situation
- A professional services firm's prior engagement survey asked whether 'communication is generally good.' 78% agreed, every year, for three years.
- Observation
- Under Klarwerk's concrete-anchor version of the same construct — whether bad news reaches leadership as fast as good news — only 41% agreed.
- Interpretation
- The original question was measuring general goodwill toward the company, not the specific mechanism leadership actually needed visibility into.
- Business Implication
- Three years of prior survey data had been giving leadership false confidence in a channel that was quietly asymmetric the entire time.
- Recommended Executive Action
- Treat prior engagement-survey trend lines as a different measurement entirely — do not attempt to reconcile them with Company Radar scores as if they tracked the same thing.
Business Impact
- Cost if this signal deteriorates
- A poorly-anchored instrument produces scores that drift a few points between cycles without reflecting any real underlying change — noise mistaken for signal.
- Executive risks that increase
- Leadership risks reacting to statistical noise as if it were a genuine finding, or dismissing a genuine finding as noise.
- Decisions that become harder
- Any decision to act on a specific dimension score depends on trusting that the question behind it actually measured something concrete.
- Typically affected
- The reliability of every dimension score the instrument produces.
Where You Will See This
03The Eight Dimensions
Each of the eight dimensions targets a distinct organizational mechanism — not a mood. Together they span from how information and authority move (Leadership, Communication, Trust & Safety) through how decisions become action (Strategy & Alignment, Innovation & Adaptability, Processes & Collaboration) to what the organization actually values and can sustain (Culture & Values, Sustainable Performance).
+
The Eight Dimensions
Each of the eight dimensions targets a distinct organizational mechanism — not a mood. Together they span from how information and authority move (Leadership, Communication, Trust & Safety) through how decisions become action (Strategy & Alignment, Innovation & Adaptability, Processes & Collaboration) to what the organization actually values and can sustain (Culture & Values, Sustainable Performance).
Detailed Explanation
Leadership measures whether leadership's stated reasoning, follow-through, and judgment under pressure are consistently confirmed by what the organization actually experiences — not whether leadership is well-liked. Two of the dimension's four scale items target this most directly: whether reasoning behind decisions is visible, and whether commitments are followed through on. The remaining two items — confidence under difficulty, and decisiveness on hard tradeoffs — are broader, related signals of leadership confidence that support the same dimension without individually establishing the same specific construct. Communication measures the velocity and directional symmetry of information flow, with particular attention to whether unfavorable information travels as fast as favorable information, since this asymmetry is a well-documented precursor to leadership being blindsided by a problem that was visible internally for months.
Trust & Psychological Safety measures whether problems, dissent, and confusion surface early and from all levels of the hierarchy — not general comfort or likability. This dimension carries the heaviest single weight of any dimension in the entire model (45% of the Risk Index) because early disclosure is the mechanism that determines whether a problem is cheap to fix or expensive.
Strategy & Alignment measures whether stated priorities are specifically understood, resourced, and consistently translated into daily decisions — not whether employees can recite the mission statement. Innovation & Adaptability combines several related signals rather than one single construct: merit-based idea evaluation, proactive versus crisis-forced adaptation, approval-stage friction, and perceived responsiveness relative to competitors. The last of these is a perception of competitive outcome, not a direct measurement of an organization's underlying capacity to reconfigure — a distinction worth holding onto when reading this dimension, since the two can diverge. Processes & Collaboration similarly combines handoff-specific signals (smoothness, ownership clarity, rework) with a broader item on whether processes generally help or hinder getting work done — the former targets coordination friction at specific team boundaries directly; the latter is a wider process-quality perception that can move for reasons beyond boundary friction alone. Processes & Collaboration measures whether documented process matches lived practice, and whether cross-team handoffs and agreements survive execution intact — a distinct failure mode from strategic misalignment, since a team can fully understand the strategy and still lose the work in translation between departments.
Culture & Values measures whether a stated value holds up when unobserved, under real conflict with a deadline, and under enforcement against a high performer specifically — the case most likely to be quietly excused. Sustainable Performance measures whether current output is being earned or borrowed against a future cost: skipped steps, concentrated dependency on specific individuals, absent recovery after demanding periods, or perpetually deferred foundational work.
Each dimension is scored from its four scale questions, averaged and normalized onto a 0–100 scale, with a minimum of five responses required before that dimension is reported for any given segment. No dimension is scored from fewer than four distinct questions, and no two questions within a dimension were retained if review found them measuring the same underlying mechanism — an internal design pass explicitly removed items that, on inspection, turned out to be near-duplicates of another question in the same dimension.
In the Dashboard & Report
The executive report's Page 2 ('Supporting Data') lists all eight dimension scores in a single table alongside their response counts — not narrated individually, but shown as the disclosed evidence underneath the three headline indices, so a reader can trace any index score back to the specific dimensions that produced it.
Limitations & Confidence
Eight dimensions cannot capture every axis of organizational health, and the choice of these eight — rather than a different set of eight — reflects a specific model of what predicts risk, execution, and sustainability, not an exhaustive taxonomy of everything that could matter about an organization. A condition genuinely affecting a company that doesn't map cleanly onto one of these eight mechanisms (for example, a highly specific supply-chain dependency) would not be captured by this instrument at all. Culture & Values is currently a supporting Company Radar dimension, grounded in its own survey-item rationale (drawing on the espoused-versus-enacted values distinction in organizational research), rather than one of the flagship Journal research essays that ground the other seven dimensions at length. This is a documentation difference, not a scoring or survey difference — the dimension is measured and weighted identically to the rest.
At a Glance
Why Executives Should Care
Each dimension maps to a specific commercial exposure, not an abstract HR category: Trust & Safety maps to how early a costly problem gets caught; Strategy & Alignment maps to whether capital and headcount decisions actually land where leadership intends; Sustainable Performance maps to whether this quarter's numbers are a reliable forecast input or a number being borrowed from next quarter. Watching these eight scores over time is a direct read on execution speed, transformation risk, and retention exposure — well before any of those show up in a financial statement.
Executive Case Study
- Situation
- A mid-size manufacturer scores Leadership 74 and Strategy & Alignment 58 in the same cycle.
- Observation
- Leadership's follow-through and judgment are strongly trusted, but the specific score for whether daily work reflects stated priorities is materially lower.
- Interpretation
- Employees trust leadership but do not consistently translate strategy into day-to-day priorities — a translation gap between the top of the organization and the work itself, not a leadership credibility problem.
- Business Implication
- Transformation initiatives are likely to slow down in execution even though they were well received when announced, because trust in the message doesn't guarantee the message reaches daily decision-making intact.
- Recommended Executive Action
- Clarify ownership and decision paths for the current strategic initiative specifically, rather than investing further in leadership communication, which is not the constrained resource here.
Business Impact
- Cost if this signal deteriorates
- A dimension left unmeasured is a category of organizational risk leadership has no visibility into at all, regardless of how well the other seven perform.
- Executive risks that increase
- Execution risk, retention risk, and strategic-drift risk can each develop undetected if the dimension tracking them is not part of the model.
- Decisions that become harder
- Prioritizing which function to invest attention in next becomes a guess rather than an evidence-based call.
- Typically affected
- Execution speed, retention, and strategic alignment, each tracked by a distinct, non-overlapping dimension.
Where You Will See This
04Index Calculation
Three executive indices — Risk, Execution, and Sustainability — are calculated as fixed, disclosed weighted averages of specific dimensions. The weights are published on every report; the arithmetic is fully reproducible from the eight dimension scores alone, by anyone, without needing to trust an editorial judgment.
+
Index Calculation
Three executive indices — Risk, Execution, and Sustainability — are calculated as fixed, disclosed weighted averages of specific dimensions. The weights are published on every report; the arithmetic is fully reproducible from the eight dimension scores alone, by anyone, without needing to trust an editorial judgment.
Detailed Explanation
Risk Index answers a single question: if something is currently going wrong, would leadership find out in time to act on it? It is calculated as 45% Trust & Psychological Safety, 30% Communication, and 25% Leadership. The weighting reflects each dimension's role in the actual sequence of a problem becoming visible: a concern must first be safe to raise (Trust & Safety) before it can travel through the organization (Communication) to someone positioned to act on it, with Leadership's credibility determining how much weight that action ultimately carries.
Execution Index answers whether a correct decision, made today, will actually be carried out. It is calculated as 40% Strategy & Alignment, 35% Processes & Collaboration, and 25% Innovation & Adaptability. Strategy & Alignment carries the heaviest weight because it measures whether a decision is even understood correctly at the point of execution; Processes & Collaboration measures whether it survives the mechanical handoffs required to implement it; Innovation & Adaptability measures the residual capacity to adjust course once real conditions diverge from the original plan.
Sustainability Index answers whether current performance can be repeated next year without a correction. It is calculated as 45% Sustainable Performance, 30% Culture & Values, and 25% Leadership. Sustainable Performance carries the heaviest weight as the most direct measure of borrowed versus earned output; Culture & Values contributes because a values gap under pressure is frequently the mechanism by which performance gets borrowed in the first place (corners get cut specifically when doing so contradicts a stated but unenforced value); Leadership's inclusion reflects that sustained performance under strain depends heavily on whether people trust leadership's judgment about what's actually necessary.
Leadership is the one dimension that contributes to two different indices, and this is an intentional design choice rather than an oversight: leadership functions as a different mechanism in each context, not as a repeated measurement of the same thing. In the Risk Index, leadership's contribution concerns whether problems, dissent, and adverse information can surface and be acted upon — a precondition for early detection. In the Sustainability Index, leadership's contribution concerns whether current performance is supported by credible decisions, real follow-through, and a realistic account of what current capacity can sustain. Neither use claims that leadership causes low risk or causes sustainable performance; both treat leadership credibility as one input, among several, to two genuinely different composite questions.
Each index is calculated only from the dimensions that currently have sufficient data (a minimum of five responses). If a dimension falls below that threshold, it is excluded from the index calculation for that cycle, and the index is calculated from the remaining dimensions at their relative weights — the index is never silently populated with a default or estimated value for a dimension that lacks real data. If fewer than two of the three dimensions feeding an index have sufficient data, that index is not calculated at all, and the report states this explicitly rather than presenting an unreliable number.
Beyond the three individual index scores, Klarwerk's decision layer classifies the combination of all three (each above or below a 60-point threshold) into one of eight fixed configuration states — for example, 'Risk low, Execution high, Sustainability low' maps to the state 'Running Hard Toward an Unseen Wall.' This configuration is the actual thesis of an executive report: not any single index in isolation, but how the three move together. A company can have a mediocre score on every individual index and still receive a fairly benign overall reading if the combination doesn't indicate a dangerous pattern; conversely, a company with two strong indices and one weak one can receive a more urgent reading if that specific combination has historically preceded a costly failure mode.
In the Dashboard & Report
The dashboard's Executive Briefing leads with the configuration label — for example, 'Risk low, Execution high, Sustainability low' — before showing any individual score, precisely because the combination, not any single number, is the finding a board member should walk away remembering.
Limitations & Confidence
Fixed weights are a defensible, disclosed choice — not a claim of having found the objectively correct weighting. They were set based on organizational research literature and internal reasoning about causal sequence (as described above), not derived empirically from a large dataset of Klarwerk customers, because no such dataset yet exists at meaningful scale. As real outcome data accumulates across customers over time, these weights are a candidate for future validation or adjustment — but as of this version, they should be understood as a considered starting model, not an empirically proven formula.
At a Glance
Risk Index
Why Executives Should Care
A single index number tells a board more in five seconds than eight individual dimension scores ever could — but only if the weighting behind it is disclosed and defensible, not an opaque roll-up. Fixed, published weights are what let a CFO or board member independently verify a finding instead of taking it on faith, which is the difference between an index a board acts on and one it politely ignores. This is the mechanism that turns eight separate signals into one number worth a resourcing decision.
Executive Case Study
- Situation
- A logistics company closes its first assessment with Trust & Safety 43, Communication 48, Strategy & Alignment 71, Processes & Collaboration 73.
- Observation
- Risk Index calculates to 50 (low band); Execution Index calculates to 69 (high band) — a 19-point gap between the two, the largest of any two indices in the report.
- Interpretation
- The organization executes fast on its decisions while its capacity to detect an emerging problem lags well behind — strong execution is not compensating for weak visibility, the two are independent.
- Business Implication
- Whatever is currently going wrong is more likely to be discovered late, and — because execution is strong — will be acted on and compounded quickly once it eventually surfaces.
- Recommended Executive Action
- Prioritize the Risk Index's underlying dimensions in the next 30 days specifically, rather than treating strong Execution as evidence the organization is broadly healthy.
Business Impact
- Cost if this signal deteriorates
- Without a disclosed, reproducible index, a board either has to trust an opaque score or discount it entirely — both outcomes waste the underlying data.
- Executive risks that increase
- A board member who cannot verify a number tends to discount it under pressure, precisely when a hard number is most needed.
- Decisions that become harder
- Board-level resourcing conversations move faster when every participant can independently reconstruct the index from disclosed weights, rather than relitigating its legitimacy each time.
- Typically affected
- How quickly a finding converts into a funded decision.
Where You Will See This
05Benchmark Methodology
Every report includes a real, calculated department-versus-company-average comparison. Klarwerk does not yet include a cross-company industry benchmark — that claim was deliberately excluded from this version rather than approximated without a defensible dataset behind it.
+
Benchmark Methodology
Every report includes a real, calculated department-versus-company-average comparison. Klarwerk does not yet include a cross-company industry benchmark — that claim was deliberately excluded from this version rather than approximated without a defensible dataset behind it.
Detailed Explanation
The benchmark that appears in every Klarwerk report today is internal: each department's aggregate score is compared against the average of all departments within that same company, for that same assessment cycle. A department is only included in this comparison once it independently clears the five-response minimum. If a department's score deviates from the company average by 15 points or more, it is flagged as a structural outlier — the single most concrete, locatable finding many reports produce, since it converts an organization-wide pattern into a specific, addressable starting point.
This calculation is fully internal to the company being assessed. It requires no external dataset, no comparison population, and no assumption about what a 'typical' company in a given industry or size class looks like — it only requires that the company being measured has more than one department reporting, which is true for the overwhelming majority of Klarwerk's target customer range of 20 to 500 employees.
An earlier design draft of this product included language suggesting comparison against 'the anonymized average of comparable organizations in the same size class' — a genuine industry benchmark. That claim was removed prior to this version, deliberately, because no dataset exists yet with enough real customers, of sufficient diversity and volume, to construct a benchmark that could survive scrutiny from a sophisticated buyer asking a simple, fair question: benchmark of what, calculated how, from how many companies? Presenting an industry-average number without being able to answer that question convincingly would have been a serious credibility risk to a product whose entire positioning rests on methodological rigor.
The underlying data structures for a future cross-company benchmark already exist in the platform's schema, unpopulated. As Klarwerk's customer base grows to a point where a benchmark could be constructed honestly — with a disclosed sample size, a disclosed methodology, and enough diversity across industries and company sizes to be meaningful — that capability can be added without any change to the survey instrument or the scoring model that produces the underlying dimension scores. It is a data-maturity question, not an architectural one.
In the Dashboard & Report
The report's 'Organizational Signal' page states a department's exact deviation from the company average in points — for example, 'Operations scores 18 points below the company average (62)' — alongside the explicit disclosure that departments are flagged at a 15-point deviation threshold, so the reader can independently judge whether the finding clears a meaningful bar.
Limitations & Confidence
Without a cross-company benchmark, a reader cannot yet know whether a given score is strong, weak, or typical relative to other organizations — only whether it is strong or weak relative to that same company's other departments. A company with uniformly low scores across every department would show no outliers under this methodology, even if every department were, by some external standard, underperforming. This is a real, disclosed gap in the current version, not a hidden one.
At a Glance
Company average: 62 — Operations flagged (deviation: -18)
Why Executives Should Care
An organization-wide average can hide a problem that is actually concentrated and fixable in one function — department-level comparison is what converts 'something seems off' into a specific, resourced starting point a leader can act on Monday morning instead of launching a company-wide initiative that spends effort on functions that were never the problem. This is a direct execution-speed and resource-allocation decision, not a reporting nicety.
Executive Case Study
- Situation
- A 40-person logistics firm's company-wide scores look moderate across the board — nothing in the critical range.
- Observation
- Operations, the largest department, scores 44 against a company average of 62 — an 18-point deviation, clearing the 15-point outlier threshold.
- Interpretation
- The moderate company-wide picture was masking a concentrated, severe condition in one specific function, diluted by three other departments scoring well above average.
- Business Implication
- A company-wide response (a new all-hands policy, a general communication push) would spend effort on three departments that don't need it and fail to specifically address the one that does.
- Recommended Executive Action
- Direct the next 30 days of investigation specifically at Operations, rather than launching an organization-wide initiative based on the aggregate score alone.
Business Impact
- Cost if this signal deteriorates
- Without department-level comparison, a concentrated problem stays diluted inside a moderate company-wide average until it is already expensive.
- Executive risks that increase
- Leadership risks funding an organization-wide initiative that spends effort on functions that were never the actual problem.
- Decisions that become harder
- Where to direct the next 30 days of attention becomes specific and defensible instead of a generalized, company-wide guess.
- Typically affected
- Speed and precision of resource allocation once a risk is identified.
Where You Will See This
06Interpretation Logic
The Executive Finding on every report is generated from a fixed, disclosed mapping of the eight possible Risk/Execution/Sustainability configurations to a specific interpretation — never freehand commentary written about a specific company. This is what makes the finding reproducible and defensible, not a matter of trusting an unseen analyst's judgment.
+
Interpretation Logic
The Executive Finding on every report is generated from a fixed, disclosed mapping of the eight possible Risk/Execution/Sustainability configurations to a specific interpretation — never freehand commentary written about a specific company. This is what makes the finding reproducible and defensible, not a matter of trusting an unseen analyst's judgment.
Detailed Explanation
Once each index is classified as high or low against the 60-point threshold, the resulting combination of three binary values produces exactly one of eight possible configurations. Each configuration maps to a fixed title, interpretation, common-mistake warning, urgency level, and boardroom question — content that is identical for every company that lands in that configuration, because the configuration itself, not any company-specific detail, is what determines the finding. A company scoring Risk 51 / Execution 68 / Sustainability 55 receives the same 'Running Hard Toward an Unseen Wall' interpretation as a company scoring Risk 49 / Execution 71 / Sustainability 52 — both fall into the same low/high/low configuration, and the underlying claim about what that combination means does not depend on the exact decimal values.
This is a deliberate design choice with a specific trade-off. The alternative — generating a unique, freehand narrative for every company based on its specific numbers — would feel more individually tailored, but would also be unfalsifiable and unreproducible: there would be no way for a reader to verify that the finding follows from a fixed, disclosed rule rather than from an unseen analyst's editorial judgment on that particular day. The configuration-based approach sacrifices some surface-level customization in exchange for a finding a board member can independently reconstruct from the three index scores and this published mapping, without taking Klarwerk's word for it.
Urgency level is calculated from the base configuration and then escalated by one level if the Sustainability trend is declining across consecutive cycles — meaning a company's second or later assessment can produce a more urgent reading than its first, purely from the trajectory, even if the underlying configuration hasn't changed. This is one of the only places in the interpretation layer where anything beyond the current cycle's raw configuration affects the output.
If confidence is below threshold (response rate under 60%) or a department-level outlier was detected, the executive recommendation text is extended with an explicit qualifier or a department-investigation instruction, appended to — not replacing — the base configuration's fixed recommendation. This ensures a low-confidence report is never presented with the same unqualified authority as a high-confidence one, without requiring an entirely separate interpretation model for low-confidence cases.
In the Dashboard & Report
Page 3 of the executive report ('Executive Finding') states the configuration explicitly — 'Risk low, Execution high, Sustainability low' — directly beneath the eyebrow, with a footnote clarifying that 'this finding reflects one of a fixed set of interpretations, selected by which indices are above or below the reporting threshold — not written specifically for this organization.' That disclosure is intentional and permanent, not a hedge added defensively.
Limitations & Confidence
A fixed, eight-state interpretation model necessarily compresses a wide range of specific numeric outcomes into a small number of narrative buckets. Two companies with meaningfully different underlying dynamics, but the same high/low classification on all three indices, will currently receive the same interpretive language — a genuine limitation this version accepts in exchange for reproducibility, and a leading candidate for refinement as more company-specific evidence (such as the specific department outlier, where one exists) becomes available to differentiate the generated text further.
At a Glance
Why Executives Should Care
A board doesn't act on eight numbers — it acts on one sentence it can repeat in the next meeting. A fixed, disclosed interpretation model is what makes that sentence defensible under questioning instead of sounding like an unverifiable editorial opinion. This directly affects how fast a finding converts into a resourcing decision: a defensible one-sentence thesis moves through a leadership team in one meeting; an opaque one gets relitigated for a quarter.
Executive Case Study
- Situation
- Two different companies both classify as 'Risk: low, Execution: high, Sustainability: low' in the same reporting cycle.
- Observation
- Their exact index scores differ by several points each, but both fall into the identical configuration bucket and receive the identical fixed interpretation, 'Running Hard Toward an Unseen Wall.'
- Interpretation
- The finding is reproducible by design — any board member can verify, from the three published scores alone, that this specific interpretation was the correct one to apply, without needing to trust an analyst's individual judgment call.
- Business Implication
- Leadership can defend the finding to a skeptical board member in real time, using only the disclosed configuration table, rather than falling back on 'that's what the report says.'
- Recommended Executive Action
- When presenting the finding, lead with the configuration label itself ('Risk low, Execution high, Sustainability low'), not just the narrative title — it's the part a board member can independently check.
Business Impact
- Cost if this signal deteriorates
- An unverifiable finding gets relitigated in every meeting it's presented in, burning leadership time that a reproducible one would not.
- Executive risks that increase
- A board that cannot independently check a finding tends to treat it as opinion, which weakens the case for acting on it.
- Decisions that become harder
- The finding is defensible in real time under direct board questioning, without needing to appeal to an unseen analyst's judgment.
- Typically affected
- How much boardroom time a finding costs before it converts into a decision.
Where You Will See This
07Open-Text Processing
Free-text responses are collected for eight questions, one per dimension, and are never displayed verbatim to anyone — including Klarwerk's own staff and the customer's leadership. They exist solely to be aggregated into themes once five or more responses to the same question exist.
+
Open-Text Processing
Free-text responses are collected for eight questions, one per dimension, and are never displayed verbatim to anyone — including Klarwerk's own staff and the customer's leadership. They exist solely to be aggregated into themes once five or more responses to the same question exist.
Detailed Explanation
Each open-ended question is deliberately worded to discourage identifying information: participants are explicitly instructed not to name any person, team, or specific event, and the question wording itself (for example, 'describe, in general terms, the kind of issue someone here would think twice before raising') is built to elicit a category of concern rather than a specific incident. This is a design-time safeguard, applied before a single response is ever collected, not a filter applied afterward.
Open-text responses are stored in a database table structurally identical in its anonymity guarantee to the scale-question responses: no field links a stored comment to the participant who wrote it. The same architecture described in the Anonymity Architecture section applies here without modification — there is no separate, weaker anonymity standard for free text.
As with scale questions, open-text responses are only surfaced in aggregate, and only once five or more responses to that specific question exist within the reporting segment. Below that threshold, open-text content for that question and segment is withheld entirely from any report or dashboard view, for the same reason scale scores are withheld below threshold: five individual responses provide enough cover that no single comment can be confidently attributed to any one person, even by someone who knows the department well.
Klarwerk V1 does not perform automated thematic clustering or sentiment analysis on open-text content. This is a deliberate scope decision, not an oversight: building a defensible thematic-summarization capability requires the same rigor this entire methodology insists on everywhere else, and a premature or unreliable summarization layer would risk misrepresenting what respondents actually said — a far worse outcome than simply not yet offering the capability. In this version, open-text responses inform Klarwerk's own qualitative understanding of a report during any human review that occurs, but are not yet algorithmically synthesized into report content shown to the customer.
In the Dashboard & Report
The survey interface itself states this policy directly to participants at the point of answering: 'Individual answers are never shown — only used for aggregated thematic analysis,' displayed as placeholder text inside every open-text field, so the promise is visible at the exact moment it matters most to the person deciding how candid to be.
Limitations & Confidence
Because Klarwerk V1 does not yet algorithmically summarize open-text content into the delivered report, the richest, most specific evidence collected by the instrument (the actual words respondents chose) is not yet directly reflected in what a customer receives — the eight open questions currently serve as a listening mechanism and a future data asset more than as a direct input to this version's deliverable. This is an explicit, disclosed scope limitation, not a claim that the content doesn't exist or isn't valuable.
At a Glance
Recurring theme identified across 5+ responses — no individual text ever shown
Why Executives Should Care
Closed-scale questions tell you a score moved; open text tells you what specifically people are actually experiencing, which is what turns a number into an executive action a leader can name concretely rather than guess at. The commercial risk of getting this wrong is real: if participants suspect a comment could be traced back to them, the richest, most specific evidence in the entire instrument disappears first — before scale scores are affected at all.
Executive Case Study
- Situation
- Five separate respondents in the Sustainable Performance open question independently mention being staffed on more than one full engagement simultaneously.
- Observation
- No individual comment is displayed anywhere, but the recurrence across five independent responses clears the aggregation threshold and is recognized as a pattern.
- Interpretation
- Concentrated overload is not an isolated complaint — it's a recurring, independently-corroborated condition, which is a materially stronger signal than any single comment could provide on its own.
- Business Implication
- This specific pattern (dual-engagement staffing) is concrete enough to investigate directly, rather than requiring leadership to interpret a vague low Sustainable Performance score without knowing what form the strain actually takes.
- Recommended Executive Action
- Review current staffing allocation for dual-engagement assignments specifically, using this as the starting hypothesis rather than a general workload review.
Business Impact
- Cost if this signal deteriorates
- If participants suspect a comment could be traced back to them, the most specific evidence in the entire instrument disappears first — before any scale score is affected.
- Executive risks that increase
- Leadership loses the one channel capable of surfacing the specific texture of a problem, not just its existence.
- Decisions that become harder
- Investigating a low score without corroborating open-text context means guessing at what form the underlying issue actually takes.
- Typically affected
- The specificity available once a pattern is identified — currently a listening mechanism, not yet a report input (see limitations).
Where You Will See This
08Confidence Scoring
Every report states a confidence level — Low, Medium, or High — derived directly from the number of participants who completed the assessment, not from any qualitative judgment. Below a minimum response count, no report is generated at all.
+
Confidence Scoring
Every report states a confidence level — Low, Medium, or High — derived directly from the number of participants who completed the assessment, not from any qualitative judgment. Below a minimum response count, no report is generated at all.
Detailed Explanation
Confidence is banded by the count of participants who completed the full assessment for that cycle: Low for 5 to 19 completed responses, Medium for 20 to 49, and High for 50 or more. These thresholds apply to the total completed-participant count for the panel, not to the number of individual answers collected — an important distinction, since each completed participant generates 32 scale answers, and counting raw answer rows rather than distinct people would produce a badly inflated and misleading confidence signal.
A minimum of five completed participants is required before any report can be generated at all, for the same reason the five-response threshold governs every other segment-level disclosure in this system: below that floor, aggregation cannot meaningfully protect anonymity, and a report would either have to withhold nearly everything or risk exposing individual responses. If a panel closes with fewer than five completed responses, no report is produced, and the customer receives an explicit notice stating this rather than a report generated on insufficient data.
Separately from the confidence band, response rate — completed participants divided by total invited — feeds directly into the interpretation layer. If response rate falls below 60%, the executive recommendation text is prepended with an explicit confidence warning before any other content, instructing the reader to treat the findings as directional rather than decisive. This is a stricter, more consequential check than the confidence band alone: a company could invite 200 people and receive 55 completed responses (comfortably into the High confidence band by absolute count) while still falling below the 60% response-rate threshold, in which case the qualifying warning still applies. Both signals are calculated and applied independently.
Confidence level and response rate are stated on the report's cover page before any finding is presented, and restated in the Executive Finding's confidence statement specifically tied to that page's claim — the same discipline of stating limitations at the point of use, not only in a general methodology appendix, that governs every other section of this page.
In the Dashboard & Report
The dashboard header displays the exact figures — for example, '34 responses · Medium confidence' — directly beneath the report title, before a reader encounters a single score, so confidence is established as context for everything that follows rather than disclosed only in fine print at the end.
Limitations & Confidence
Confidence banding is a proxy for reliability, not a guarantee of it. A High-confidence report (50 or more responses) with a highly skewed department composition — for example, one department heavily overrepresented relative to its actual headcount share — could still produce a misleading company-wide picture despite clearing the absolute response-count threshold, since the confidence calculation does not currently weight for proportional representation across departments.
At a Glance
Low
5–19 responses
Medium
20–49 responses
High
50+ responses
Why Executives Should Care
A finding presented with false confidence is more dangerous than no finding at all — it can drive a resourcing decision on data that doesn't actually support it. Explicit confidence bands are what let a leadership team calibrate how much weight to put behind a result, which directly protects decision quality: a Low-confidence signal should prompt a follow-up conversation, not a reorg.
Executive Case Study
- Situation
- A 40-person company's first cycle closes with 34 completed responses (85% response rate).
- Observation
- This clears both the 60% response-rate threshold and the 20-response Medium-confidence band, with no confidence warning attached to the executive finding.
- Interpretation
- The report can be treated as directly decision-grade — contrast with a hypothetical second company closing at 12 of 40 (30%), which would fail the response-rate check entirely and carry an explicit warning prepended to every recommendation.
- Business Implication
- Leadership can act on this specific report without first commissioning a second data-gathering exercise to validate it.
- Recommended Executive Action
- Note the confidence level explicitly when forwarding the report to the board — a High or Medium confidence label is itself part of the finding's credibility, not just a footnote.
Business Impact
- Cost if this signal deteriorates
- Acting on a low-confidence finding as if it were high-confidence risks a resourcing decision the underlying data cannot actually support.
- Executive risks that increase
- A false-confidence decision is harder to reverse than no decision at all, since it consumes budget and credibility before the gap is discovered.
- Decisions that become harder
- Whether a finding warrants immediate action or a follow-up conversation first depends entirely on the stated confidence band.
- Typically affected
- How aggressively leadership can act on any single finding.
Where You Will See This
09Methodological Limitations
This report identifies patterns and conditions, never causes. It never assigns fault to a named individual, never predicts a specific future outcome, and never generalizes findings beyond the organization assessed. These are not legal disclaimers — they are the boundary of what this kind of measurement can honestly claim.
+
Methodological Limitations
This report identifies patterns and conditions, never causes. It never assigns fault to a named individual, never predicts a specific future outcome, and never generalizes findings beyond the organization assessed. These are not legal disclaimers — they are the boundary of what this kind of measurement can honestly claim.
Detailed Explanation
Klarwerk does not establish causal relationships between dimensions. If Trust & Psychological Safety and Sustainable Performance are both low in the same report, the instrument has not demonstrated that one caused the other, or that either caused anything else — only that both conditions currently exist, measured independently, in the same organization at the same time. Any causal story connecting them would require investigation this instrument does not and cannot perform on its own; the report is explicitly designed to prompt that investigation, not to substitute for it.
Klarwerk does not predict future outcomes. A low Sustainability Index describes current conditions consistent with performance being borrowed against a future cost — it is not a forecast that a specific event (attrition, a quality failure, a missed target) will occur by a specific date. The language used throughout every report is deliberately calibrated to this limit: 'may indicate,' 'is consistent with,' and 'suggests' appear where a lesser methodology might claim certainty it hasn't earned.
Klarwerk makes no claim about any individual, and this is enforced structurally, not just rhetorically: the five-response minimum makes it mathematically impossible to report a result attributable to a specific person, and no report, dashboard, or export at any tier ever displays an individual-level score. A department outlier identifies a location worth investigating — never a person to blame, and never evidence sufficient on its own to justify a personnel decision.
Klarwerk does not generalize findings beyond the specific organization assessed. A finding about one company's Meridian-style 'Running Hard Toward an Unseen Wall' pattern says nothing evidentiary about any other company, including one in the same industry or of similar size, until that other company has been independently assessed. The fixed instrument and fixed scoring model create comparability across cycles for the same company; they do not create a license to extrapolate one company's result onto another's.
Finally, Klarwerk's scoring model — the eight dimensions, the three index weights, the eight interpretation configurations — reflects a considered, disclosed design built on established organizational psychology and management research, not an empirically validated formula proven against large-scale outcome data specific to this instrument. This page states that distinction plainly rather than implying a level of empirical certainty the model has not yet earned. As the customer base and cycle history grow, aspects of this model are expected to be refined against real outcomes — and any such refinement will be disclosed here, not made silently.
In the Dashboard & Report
The final page of every executive report is titled 'Methodology & Limitations' and opens with the same sentence that governs this entire page: 'This report measures current organizational conditions using a fixed, repeatable method. It does not diagnose causes, and it does not forecast outcomes.' It is placed last, deliberately, so it governs everything the reader has just read, rather than being skimmed as a preamble before the findings arrive.
Limitations & Confidence
This section is, itself, subject to a limitation worth naming: stating a boundary clearly does not guarantee a reader will observe it in practice. A CEO who reads every caveat on this page can still, in the moment of a difficult decision, be tempted to treat a pattern as a cause or a department flag as a verdict on a specific manager. Klarwerk's Executive Readiness materials exist specifically to reduce that risk before an assessment ever launches — but no amount of disclosure fully substitutes for a customer's own discipline in applying a diagnostic instrument as what it is, not as more than it claims to be.
At a Glance
Why Executives Should Care
Overclaiming what a diagnostic instrument can determine is how organizations end up making a personnel decision on data that was never designed to support one — a governance and legal exposure, not just a methodological nuance. Being explicit about the boundary is what allows leadership to use the finding aggressively where it's strong (locating a pattern) while staying disciplined where it isn't (assigning a cause), which is exactly the distinction that protects both decision quality and the organization's legal footing.
Executive Case Study
- Situation
- A CEO reads a report showing Trust & Psychological Safety scoring 20 points below every other dimension in one specific department.
- Observation
- His first instinct is to ask HR to review that department's manager's performance file.
- Interpretation
- The finding identifies a pattern and a location — it does not identify a cause, and the manager's individual performance was never measured by this instrument at any point.
- Business Implication
- Acting on the score as if it were evidence against a specific person would be methodologically unsupported, and would teach the organization that honest participation eventually leads to someone being blamed — degrading every future cycle's honesty.
- Recommended Executive Action
- Use the finding to open a direct, non-attributive investigation into that department's conditions — not a review of any individual's file.
Business Impact
- Cost if this signal deteriorates
- Treating a pattern as a cause, or a location as a verdict on a person, is how a diagnostic tool becomes a governance and legal liability instead of an asset.
- Executive risks that increase
- Personnel action taken on data this instrument was never built to support is difficult to defend after the fact, to a board or to a court.
- Decisions that become harder
- Knowing exactly what the report does and does not claim determines how aggressively leadership can act on a finding without overreaching.
- Typically affected
- The organization's legal and governance exposure when a finding is acted upon.
Where You Will See This
10What Klarwerk Cannot Determine
The goal of this section is not to sound impressive. It's to state the boundary of this instrument as plainly as everything else on this page states its capability — because a diagnostic tool that won't name its own limits isn't one a board should trust with a resourcing decision.
+
What Klarwerk Cannot Determine
The goal of this section is not to sound impressive. It's to state the boundary of this instrument as plainly as everything else on this page states its capability — because a diagnostic tool that won't name its own limits isn't one a board should trust with a resourcing decision.
Klarwerk cannot determine
- Root causes behind any finding
- Individual responsibility for any result
- Future events or specific outcomes
- Employee intent behind an answer
- The management quality of any specific person
Klarwerk can determine
- Organizational patterns across dimensions
- Hidden contradictions between stated and lived reality
- Communication and information-flow risks
- Trust and psychological-safety gaps
- Strategic alignment problems
- Emerging risks before they become visible elsewhere
This is not a legal disclaimer appended for caution. It is the actual, honest boundary of what a structured, anonymized measurement instrument can responsibly claim to know about an organization — stated once, clearly, so it never needs to be inferred from hedged language elsewhere on this page.
11Why Organizations Measure Repeatedly
One assessment establishes a baseline. It cannot, on its own, distinguish a genuine structural condition from a temporary fluctuation caused by a single hard quarter, a recent reorganization, or ordinary variance in who happened to respond.
+
Why Organizations Measure Repeatedly
One assessment establishes a baseline. It cannot, on its own, distinguish a genuine structural condition from a temporary fluctuation caused by a single hard quarter, a recent reorganization, or ordinary variance in who happened to respond.
A single assessment answers one question: what does this organization look like right now? That is a real and useful answer on its own — it is the entire basis of everything described elsewhere on this page. But current-state measurement has an inherent limit: it cannot tell you whether what it found is a persistent condition of the organization or a temporary artifact of the specific month it happened to be measured in.
A second assessment changes what can be known. Comparing two cycles reveals direction — whether a specific dimension moved, and by how much. A score that holds steady between two cycles is meaningfully different from one that moved 15 points in either direction, and neither of those facts is available from a single measurement in isolation.
A third assessment, and every one after it, changes what can be known again. Two data points can show direction, but they cannot distinguish a genuine trend from a single fluctuation — a score could move once for a reason unrelated to any underlying condition, and a two-point comparison has no way to tell the difference. Three or more cycles begin to reveal organizational development: whether a change is sustained, whether an intervention actually worked, and whether a pattern that looked concerning in isolation was in fact part of a normal range for that specific organization.
This is why repeated measurement increases decision confidence in a way that no single assessment, however well-designed, can replicate on its own. It is not an argument for measuring more often for its own sake — it is a direct consequence of what a single point of data can and cannot distinguish. An organization that measures once knows its current state. An organization that measures repeatedly knows its trajectory, which is very often the more decision-relevant fact.
Executive Closing Note
Klarwerk does not replace leadership. It reduces uncertainty.
This methodology exists to convert weak, informal organizational signals — the kind every leadership team senses but rarely has evidence for — into structured, disclosed, reproducible executive intelligence. The report does not make a decision. It gives leadership a clearer basis on which to make one.
Ready for your first Company Radar
Organizational Intelligence emerges when every perspective matters.
Bring your organization in, invite your team, and receive strategic intelligence leadership can act on — fully anonymous, within 48 hours.