Klarwerk Insights
← Back to the Journal

Trust & Psychological Safety · Flagship Research

Psychological Safety Is a Business Variable

The term sounds soft. The research is not — and most leadership teams have never seen a rigorous, honest account of what it actually shows, where it stops, and what an organization can do with it.

January 10, 2026 · 37 min · Fully sourced, see References

A Problem Leadership Sees Only Half Of

A senior engineer notices, three weeks before launch, that a load-bearing assumption in the project plan doesn't hold. She has raised something like this before. Last time, the response was polite, mildly dismissive, and the meeting moved on. This time she says nothing, flags a minor cosmetic issue instead, and lets the deadline arrive.

Six weeks later, the assumption fails in production. The postmortem identifies a technical root cause, assigns a remediation owner, and closes the ticket. What it does not identify — because nothing in the postmortem process is built to see it — is that someone already knew, three weeks earlier, and chose not to say so.

This is not a story about one disengaged employee. It is the ordinary, statistically unremarkable behavior of a competent person operating inside a specific set of conditions. Multiply it across a few hundred employees and a few years, and it becomes one of the most consequential and least visible sources of organizational risk a CEO will ever have to manage — precisely because, by construction, it never shows up as a single dramatic event. It shows up as a pattern of things leadership found out about later than it should have.

The research literature has a name for the condition that determines whether the engineer speaks or stays quiet: psychological safety. It is one of the most extensively studied constructs in organizational behavior, and one of the most casually misunderstood in ordinary management conversation — frequently reduced to a synonym for niceness, comfort, or the absence of conflict. None of those are accurate. Understanding what the research actually supports, and where it stops, is the difference between a leadership team that treats this as a soft HR concern and one that treats it as what the evidence indicates it actually is: a measurable precondition for how early an organization finds out about its own problems.

Why This Belongs on a CEO's Agenda, Not Only HR's

The immediate objection is reasonable: this sounds like a culture topic, and culture topics are notoriously difficult to connect to anything a CEO is actually accountable for. The connection here, however, is not sentimental. It runs through four specific executive concerns: decision quality, error detection, execution speed, and adaptability under changing conditions — each of which has a body of research behind it, not just an assertion.

Decision quality depends on the information available at the point of decision. A leadership team makes better decisions when the full range of relevant concerns, including inconvenient ones, actually reaches the table before the decision is made rather than after. Psychological safety research is fundamentally about what determines whether that happens — whether a specific piece of information travels from the person who has it to the person who needs it, or whether it stays with the person who has it because raising it carries a cost they'd rather not pay.

Error detection is a related but distinct concern. Every organization makes mistakes; the question that actually differentiates organizations is how quickly a mistake is noticed, reported, and corrected once it occurs. Amy Edmondson's earliest empirical work on this construct was conducted specifically to understand variation in error detection across hospital nursing units, and found that teams differed not primarily in how many errors occurred, but in how likely those errors were to be surfaced and addressed once they had (Edmondson, 1996). That distinction — error rate versus error detection rate — turns out to matter enormously for any organization, not just hospitals, because an error that is caught early is inexpensive, and the same error caught late is not.

Execution speed is affected in a less obvious way. A commitment to move fast is only as good as the organization's ability to notice, in real time, that a specific execution path isn't working. An organization where that kind of course-correcting information travels freely can afford to move quickly, because it will find out fast if something needs to change. An organization where it doesn't has to either move more cautiously or accept that it is running blind on stretches of the path — a tradeoff that is rarely made consciously, because leadership often does not know which condition it is actually operating under.

Adaptability, the capacity to respond to a change the organization did not initiate, depends on some version of the same mechanism: does information about a shifting external condition reach the people positioned to act on it, or does it stall somewhere in the hierarchy because raising it would mean contradicting a plan someone senior has already committed to? The research literature on this specific link is younger and thinner than the literature on learning behavior or error detection, and this article will be explicit later about exactly how much weight it can bear.

None of this requires speculative financial modeling to matter to a CEO. It requires taking seriously that a specific, well-documented psychological condition determines whether the information a leadership team needs to make good decisions, catch expensive mistakes early, and move quickly with confidence actually reaches them — and that most leadership teams have no direct, systematic way of knowing where they currently stand on it.

What the Term Actually Means — And What It Doesn't

The most-cited definition comes directly from the researcher who did more than anyone to establish the construct empirically. Amy Edmondson defined team psychological safety as "a shared belief held by members of a team that the team is safe for interpersonal risk taking" (Edmondson, 1999). Every word in that definition is doing specific work, and most casual usage of the term drops several of them.

It is shared, not individual. Psychological safety in this literature is a property of a team or group — a collective perception, not simply one person's private comfort level. Two people on the same team can, in principle, have somewhat different individual perceptions, but the construct as studied is about the climate the group has collectively established, not a personality trait of any one member.

It concerns interpersonal risk specifically — the risk of looking ignorant by asking a question, looking incompetent by admitting a mistake, looking negative by raising a concern, or looking disruptive by challenging the status quo. This is a narrower and more precise target than general workplace comfort. A team can be low-stress, friendly, and pleasant to work on while still carrying real interpersonal risk around these four specific behaviors — and a team can carry real tension and disagreement while still being genuinely safe for people to admit what they don't know.

That distinction is the single most important thing to get right, because it is also where the term is most often misapplied in ordinary management conversation. Psychological safety is not the same as psychological comfort. A team can be uncomfortable — working on a hard problem, under real deadline pressure, disagreeing sharply about the right approach — and still be psychologically safe, if the discomfort is about the work rather than about the personal risk of speaking honestly. Comfort and safety can move independently of each other, and conflating them leads to exactly the wrong management response: softening the work itself instead of addressing whether people can be honest about it.

It is also not the same as interpersonal trust, although the two are related and often travel together. Trust is typically framed as a belief about another specific person's reliability or benevolence. Psychological safety is a belief about the group's climate — what will happen, generally, if you take an interpersonal risk in this particular room. You can trust a specific colleague completely and still feel it isn't safe to challenge a decision in a meeting where three more senior people are present. The climate, not just the individual relationship, is what the construct is measuring.

It is not the same as liking your colleagues, and it is not consensus. A team with strong psychological safety is not necessarily a team where everyone agrees or gets along easily — in fact, one of the behaviors the construct is meant to protect is exactly the opposite: the willingness to disagree, challenge an idea, or say something the group doesn't want to hear. A team that suppresses disagreement in the name of harmony is not psychologically safe by this definition; it is conflict-avoidant, which is a related but distinct and in some ways opposite pattern.

Perhaps most importantly for an executive audience: psychological safety is not the absence of accountability, and it is not low standards. Edmondson has been explicit and consistent on this point across her body of work, distinguishing psychological safety from what she calls a permissive or "anything goes" climate. The clearest formulation treats psychological safety and high standards as two separate dimensions that combine to produce different organizational climates — high safety with low standards produces a comfortable but underperforming team; high standards with low safety produces an anxious, defensive team that hides problems; and it is the combination of both that the research associates with genuine learning and performance. Safety is what makes it possible to hold high standards honestly, by ensuring that falling short of them can be discussed rather than concealed.

Getting this definition precisely right matters because every misapplication in the list above leads to a different, and usually wrong, management intervention. If a CEO believes psychological safety means comfort, the intervention becomes making work less stressful, which does nothing for the underlying issue. If a CEO believes it means consensus, the intervention becomes suppressing productive disagreement, which actively damages the thing the construct is meant to protect. The precise definition is not academic pedantry — it is the only version of the concept that actually points toward the right organizational lever.

The History Behind the Construct

Psychological safety did not originate as a management buzzword; it has a specific, traceable intellectual lineage inside organizational psychology, and understanding that lineage helps explain why the construct is taken as seriously as it is in the research literature.

An early and influential precursor comes from William Kahn's 1990 study of personal engagement and disengagement at work, published in the Academy of Management Journal. Kahn examined the psychological conditions under which people brought their full selves — physically, cognitively, and emotionally — into a work role versus withdrawing and disengaging, and identified psychological safety as one of three core conditions (alongside meaningfulness and availability) that determined which of the two happened (Kahn, 1990). This work established the conceptual groundwork: safety as a precondition for genuine engagement, rather than a nice-to-have adjacent to it.

The construct's decisive empirical development, and the version most subsequent research builds on, came from Amy Edmondson's work through the 1990s. Her 1996 study, conducted across nursing units in two hospitals, was designed to investigate variation in medication error rates — and produced a genuinely counterintuitive finding that became foundational to the entire field. The study found systematic differences across units not only in how many errors occurred, but also in how likely an error was to be detected, discussed, and learned from once it had (Edmondson, 1996). That second finding is the important and frequently under-appreciated one: it means a unit's outwardly cleaner error record cannot be read at face value, since part of the gap between units may reflect differences in disclosure rather than differences in the underlying rate of mistakes. The study does not establish that any two units had identical true error rates — only that the observable, reported metric and team climate were entangled in a way that should make a leader cautious about reading a low reported-error count as straightforward good news. This distinction between what is reported and what actually occurred is one of the most frequently flattened lessons in this body of research, and it recurs, in different forms, throughout the rest of this article.

Edmondson formalized the construct three years later in her widely cited 1999 paper in Administrative Science Quarterly, introducing team psychological safety as a distinct construct, developing a validated survey measure for it, and testing it empirically across 51 work teams in a manufacturing company. That study found that team psychological safety was significantly associated with team learning behavior, and that learning behavior in turn mediated the relationship between psychological safety and team performance (Edmondson, 1999). This is the paper most subsequent psychological safety research treats as the field's foundational reference point.

The construct then moved through what Edmondson and her co-author Zhike Lei later described, in a 2014 review, as a period of relative dormancy through the early 2000s followed by a substantial "renaissance" of research activity from roughly the mid-2000s onward, as scholars extended the construct beyond its original team-level formulation to examine it at the individual, organizational, and cross-cultural levels (Edmondson & Lei, 2014). Their review is itself one of the most useful entry points into this literature for a non-specialist reader, because it explicitly maps what had been established with reasonable confidence by 2014 against what remained genuinely open.

Applied and popular interest in the concept accelerated sharply after 2015, when Google published internal findings from what it called Project Aristotle — a two-year internal study of roughly 180 of its own teams, examining what distinguished more effective teams from less effective ones. Google's People Analytics team, working from a substantial dataset of internal team characteristics, identified psychological safety as the strongest single differentiator among the team dynamics they studied. This finding was popularized widely, most notably through a New York Times Magazine feature by Charles Duhigg in February 2016, and through Google's own public re:Work materials.

It is worth being precise about what Project Aristotle is and is not, because it is simultaneously the most widely recognized touchpoint for this concept in business audiences and methodologically the weakest source in this article's reference list. It was an internal, observational corporate study, not peer-reviewed academic research; it was never subjected to external peer review, its methodology was not published with the same rigor or transparency as an academic paper, and — as an observational study of existing team differences rather than a controlled or longitudinal design — it cannot establish that psychological safety caused better performance rather than simply co-occurring with it. It is genuinely useful as an applied case demonstrating that a rigor-minded technology company, starting from a very different hypothesis (that team composition and individual talent would matter most), arrived independently at a conclusion consistent with decades of prior academic research. It should not be cited, and is not cited here, as evidence with the same standing as the peer-reviewed literature that came before it.

What the Evidence Actually Supports — and Where It Gets Thinner

The most rigorous single answer to "what does the research show" available for this article comes from a 2017 meta-analysis by Frazier, Fainshmidt, Klinger, Pezeshkan, and Vracheva, published in Personnel Psychology. The authors aggregated 136 independent samples, representing more than 22,000 individuals across nearly 5,000 groups, to synthesize what the accumulated empirical literature actually supports about the antecedents and consequences of psychological safety (Frazier et al., 2017). A meta-analysis of this scale is close to as strong a form of correlational evidence as exists in organizational behavior research, precisely because pooling findings across many separate studies and samples cancels out much of the noise, sampling error, and idiosyncrasy that limit confidence in any single study — including the antecedent-focused Frazier et al. (2017) synthesis itself distinguishes between what predicts psychological safety (its antecedents) and what psychological safety in turn predicts (its outcomes), a distinction worth holding onto rather than treating both halves of the model as a single undifferentiated finding. What a meta-analysis of this kind cannot do, no matter its scale, is convert the underlying studies' predominantly cross-sectional, correlational design into causal evidence; more data points strengthen confidence that a pattern is real and replicable, not confidence about which variable, if either, is driving the other. A separate systematic review by Newman, Donohue, and Eva (2017), published in Human Resource Management Review, reached a broadly consistent conclusion through a different method — rather than statistically pooling effect sizes, the authors mapped the accumulated empirical literature's antecedents, outcomes, and moderators of psychological safety across individual, team, and organizational levels of analysis, and explicitly flagged the same gap Frazier et al. later quantified: a literature comparatively rich in outcome studies and comparatively thin on rigorously established causes (Newman, Donohue, & Eva, 2017). Two independent reviews, using different methods, arriving at the same limitation, is itself a modestly reassuring sign that the limitation is real rather than an artifact of one research team's approach.

Their synthesis found individual-level psychological safety to be strongly associated with work engagement, job satisfaction, and organizational commitment, and moderately associated with task performance and with organizational citizenship behaviors — the discretionary, extra-role contributions like helping colleagues or proactively suggesting improvements that go beyond formal job requirements. Group-level psychological safety showed a broadly similar pattern of associations, though the authors note there was insufficient available data in the literature at that time to examine some of the same relationships (commitment and citizenship behavior specifically) at the group level with confidence.

It is important to be precise about what "strongly associated" and "moderately associated" mean here, because meta-analytic effect sizes are frequently overstated in secondary summaries. These are correlational relationships aggregated across many studies — they establish a consistent statistical pattern across a large combined sample, not proof that psychological safety is the single cause of engagement or performance in any given organization. Many other factors plausibly affect both variables simultaneously, and the majority of the underlying studies in this literature are cross-sectional, meaning safety and its outcomes were measured at the same point in time rather than safety being measured first and an outcome observed afterward.

This leads directly to the most important limitation in the entire literature, and one the meta-analysis's own authors state explicitly rather than leaving implicit: due to the scarcity of experimental research on this topic, there is limited evidence establishing what actually causes psychological safety to develop in the first place, as distinct from what it correlates with once present (Frazier et al., 2017). This is a genuinely significant gap. It means the field can say with real confidence that psychological safety and various positive outcomes travel together with notable consistency across a large combined body of research; it is on much thinner ground establishing the direction of causation, or ruling out the possibility that some third factor — a strong general manager, for instance, or a well-resourced team — drives both safety and performance independently, making them correlated without either one causing the other.

A separate and important thread within this literature concerns learning behavior specifically, distinct from performance more broadly. Edmondson's original 1999 study found that psychological safety predicted learning behavior directly, and that learning behavior — not psychological safety on its own — was the variable that most directly predicted team performance, functioning as the mechanism connecting the two (Edmondson, 1999). This is a meaningfully different claim than "psychological safety causes performance." The more precise and better-supported version is closer to: psychological safety creates the conditions under which a team is willing to engage in the specific behaviors that constitute learning — asking questions, seeking feedback, discussing errors, experimenting, and reflecting on outcomes — and it is that learning behavior, not the underlying safety belief itself, that plausibly drives improved performance over time.

Taken together, the state of the evidence supports a confident claim that psychological safety is real, measurable, consistently associated with a specific and coherent set of workplace behaviors and attitudes across a very large combined body of research, and mechanistically linked to learning behavior specifically in a way that has direct theoretical and practical relevance to how organizations catch and correct problems. It supports a much more cautious claim about causality, about what specifically produces psychological safety in a given team, and about how large its effect is relative to other organizational factors operating at the same time.

Psychological Safety and the Willingness to Speak

Of all the behaviors psychological safety research examines, the willingness to speak up — sometimes studied under the closely related label of employee voice — is the one most directly relevant to the executive concerns this article opened with: whether a leadership team finds out about a problem while it is still cheap to address.

The mechanism the research proposes is straightforward and intuitive once stated plainly: raising a concern, admitting uncertainty, disagreeing with a decision, or reporting a mistake all carry some real or perceived interpersonal risk. Whether that risk feels survivable is what psychological safety describes. Where the risk feels too high, the behavior doesn't happen — not necessarily because the person doesn't notice the problem, but because staying silent is the individually rational choice given the perceived cost of speaking.

Edmondson's 1996 hospital study remains the clearest empirical illustration of why this matters organizationally rather than just individually: it found that units differed both in error frequency and, independently, in how likely an error was to be surfaced and discussed once it occurred, meaning a unit's reported error count is not a reliable stand-in for its actual rate of mistakes (Edmondson, 1996). Later work by Nembhard and Edmondson extended this line of investigation into the role leaders specifically play, examining how leader inclusiveness — the degree to which a leader actively invites and appreciates others' input — affects psychological safety and, downstream, quality improvement efforts in healthcare teams (Nembhard & Edmondson, 2006). Carmeli and Gittell (2009), working across two separate studies outside the healthcare setting, found that psychological safety statistically mediated the relationship between the quality of working relationships among colleagues and the degree to which employees engaged in learning from failures — evidence that the error-disclosure pattern Edmondson identified in hospitals extends to a broader mechanism connecting how people relate to one another and whether failure becomes a source of organizational learning or something quietly buried (Carmeli & Gittell, 2009). Their findings connected leader behavior directly to whether staff engaged in the kind of improvement-oriented voice behavior that requires some willingness to point out that something isn't working.

An important nuance the research is careful about, and that an executive audience should be equally careful about, is that silence is not automatically evidence of a hidden problem, and every silent employee is not concealing something leadership needs to know. People stay quiet for many reasons that have nothing to do with psychological safety: the concern may already be known and being addressed elsewhere, the person may not consider it significant enough to raise, or they may simply have nothing to add in a given moment. Psychological safety research is about the systematic pattern — whether, in aggregate and over time, a team's climate makes people more or less likely to surface something significant when they do have it — not a claim that can be applied to interpret any single individual's silence in any specific instance.

This distinction matters for how the construct should and shouldn't be operationalized in an organizational measurement context. The useful question is not "is everyone speaking up all the time," which is neither a realistic nor even a desirable state — plenty of moments genuinely don't call for input. The useful question is closer to: across the organization, does the climate systematically make certain kinds of information, particularly unfavorable or uncertain information, less likely to travel than favorable and certain information? That asymmetry, not the raw amount of speaking-up behavior, is what the evidence suggests is worth an organization's sustained attention.

Psychological Safety and Performance — A Careful Reading

It would be convenient, and it is a common shorthand in business writing, to state simply that psychological safety causes high performance. The research does not support that sentence as written, and it is worth being explicit about exactly why, because the more accurate version is also the more useful one for a CEO deciding what to do with this information.

Start with what is well-established: psychological safety is consistently, and at meta-analytic scale, associated with task performance — the Frazier et al. (2017) meta-analysis characterizes this as a moderate association at the individual level, smaller than the strong associations found for engagement, satisfaction, and commitment, but real and consistent across the aggregated literature. Edmondson's original 1999 study adds an important structural detail to this picture: in her data, the pathway ran through learning behavior specifically, not directly from safety to performance. Teams with higher psychological safety engaged in more learning behavior — more questioning, more feedback-seeking, more open discussion of mistakes and near-misses — and it was that learning behavior which predicted performance, with psychological safety functioning as its precondition rather than as a direct cause of outcomes on its own.

This distinction has real practical consequence. It suggests that psychological safety on its own, without a team or organization that actually does something with the resulting openness — discusses the mistake, adjusts the process, tests the new idea — may not translate into better outcomes by itself. A team can become safer without automatically becoming more effective, if safety isn't paired with a genuine mechanism for acting on what that safety surfaces. This is a meaningfully more precise and more actionable claim than "safety causes performance," because it points toward a second condition — some organizational capacity or willingness to actually respond to what gets raised — that has to be present alongside safety for the performance benefit to materialize.

It is also important to hold performance, learning behavior, and team effectiveness as related but genuinely distinct constructs rather than treating them as interchangeable, which secondary business writing on this topic frequently does. Learning behavior refers to the specific behavioral pattern — asking, testing, reflecting, discussing failure — that safety is theorized to enable. Performance is an outcome, typically measured through some combination of output quality, goal achievement, or externally rated effectiveness. Team effectiveness is broader still, sometimes incorporating member satisfaction and the team's ongoing viability alongside its task output. Psychological safety research has the clearest and most consistent evidence connecting to learning behavior specifically; the connection to performance is real but somewhat less direct and less uniformly strong across studies; and claims about overall team effectiveness require aggregating across an even wider and more heterogeneous set of findings.

None of this should read as a reason to dismiss the performance connection — a moderate, consistent association across more than 22,000 individuals and nearly 5,000 groups is a real and meaningful finding, not a weak one. It should read as a reason to be precise about what kind of claim is actually supportable: psychological safety creates conditions favorable to learning behavior, which is itself associated with better performance, and the overall relationship between safety and performance holds up as a real, moderate, well-replicated pattern — without being reducible to a simple, one-directional causal statement a CEO could act on as if it were settled fact.

Psychological Safety, Innovation, and Adaptability — What's Supported and What Isn't

The connection between psychological safety and innovation is intuitively appealing — proposing a new idea is itself a form of interpersonal risk-taking, since a bad or unconventional idea can carry a real social cost to whoever raises it — and it appears frequently in applied and popular writing about the construct. The underlying academic evidence for this specific link is real but considerably thinner and less consolidated than the evidence base for learning behavior or error detection.

Edmondson and Lei's 2014 review discusses innovation and creativity as one of several outcome domains where psychological safety research has extended, noting a body of work connecting safety to constructs like creative involvement and idea generation at the individual and team level. Several specific studies illustrate what this extension looks like in practice, without yet amounting to a meta-analysis of the scale of Frazier et al. (2017). Baer and Frese (2003), studying process innovations and financial performance across a sample of German firms, found that a climate combining psychological safety with initiative was associated with higher process innovation and, in turn, with firm performance — one of the few studies in this literature to connect the construct to an organizational-level outcome rather than an individual or team-level one. Kessel, Kratzer, and Schultz (2012), studying healthcare teams, found psychological safety associated with knowledge-sharing behavior and, through it, with team creative performance. These are genuine, peer-reviewed empirical findings, not merely theoretical extrapolation — but they remain a smaller and more fragmented body of work than the engagement, performance, and citizenship-behavior findings Frazier et al. (2017) were able to synthesize at scale, and a CEO should weight them accordingly: real and worth taking seriously, not yet as thoroughly cross-validated as the core construct's other outcomes.

Adaptability — an organization's capacity to respond to a change it did not initiate — is connected more loosely still, resting more on theoretical extension than on a dedicated, robust empirical base specific to that outcome. The closest available empirical anchor is Edmondson, Bohmer, and Pisano's (2001) field study of sixteen hospitals implementing an identical new surgical technology, which found that team-level learning behavior — the same behavior their earlier work tied to psychological safety — was a key factor distinguishing hospitals that successfully adapted their routines to the new technology from hospitals that did not. That study is about technology adoption specifically, not organizational adaptability in the broader sense this article has been using the term, and it examines learning behavior as the mechanism rather than measuring psychological safety directly as a predictor. It is included here as the most concrete available illustration of the underlying logic, not as direct proof of it: an organization where people feel safe flagging that a plan isn't working should, in principle, be able to redirect faster than one where that information gets suppressed until the plan has already visibly failed — a claim this article states because it is coherent and consistent with the broader learning-behavior findings, not because a dedicated, robust body of research has isolated psychological safety's effect on adaptability as distinct from its effect on learning behavior more generally.

This is precisely the kind of distinction intellectual honesty requires making explicit rather than blurring for narrative convenience. The learning-behavior and error-detection findings in this article rest on a genuinely strong evidentiary foundation. The innovation and adaptability connections are real, actively studied, and theoretically well-motivated extensions of that foundation — but a CEO reading this article should understand that they currently rest on a thinner and less consolidated base of evidence than the core claims in the sections before this one.

How the Concept Gets Misused — And Why This Section Matters

Every widely adopted management concept eventually gets stretched to justify things its originators never intended, and psychological safety has been stretched in several specific, predictable, and organizationally costly directions. A CEO who finishes this article with a sharper sense of what the concept is not will be better protected against these misapplications than one who only absorbs what it is.

The first and most common distortion: psychological safety does not mean low accountability. Edmondson has consistently framed psychological safety and high performance standards as two independent dimensions of team climate, not substitutes for each other. A team that is safe but has no real standards produces what her framework describes as a comfortable, low-performing climate — pleasant, low-friction, and not particularly effective. The organizational value of psychological safety comes specifically from pairing it with real accountability, not from replacing accountability with comfort. A leader who hears "psychological safety" and responds by softening expectations or avoiding hard performance conversations has inverted the concept, not applied it.

The second distortion: psychological safety does not mean avoiding difficult conversations. If anything, the research points in the opposite direction — a genuinely safe team should be more able to have a difficult conversation, not less, because the interpersonal risk of raising something hard has been reduced rather than because the hard thing has been avoided. An organization that uses "psychological safety" as a reason to defer a necessary but uncomfortable conversation about someone's performance, or a strategic disagreement that needs resolving, has confused safety with conflict avoidance — two things this article's earlier definitional section was explicit about being different, and in some respects opposed to each other.

The third distortion: psychological safety does not mean everyone agreeing, or decisions being made by consensus. A team where disagreement has been suppressed in the name of harmony is not exhibiting high psychological safety by the research definition — it is exhibiting a specific and well-documented failure mode where the social cost of dissent has simply been made invisible rather than eliminated. Genuine psychological safety should, if anything, make disagreement more visible and more frequent in a well-functioning team, not less, because raising a dissenting view no longer carries the same interpersonal cost.

The fourth distortion: psychological safety does not mean making every employee comfortable at all times, or protecting people from the ordinary difficulty of demanding work. Difficult, ambiguous, high-stakes work is inherently uncomfortable regardless of team climate. The construct is about whether that discomfort can be discussed honestly — whether someone can say "I'm not sure this is working" or "I don't understand this requirement" without a social penalty — not about eliminating the discomfort of hard work itself.

A leadership team that internalizes these distinctions arrives at a materially more precise and more actionable understanding than one working from the popular, flattened version of the concept: psychological safety is not a program for making work pleasant. It is a specific organizational condition under which honest information about problems, uncertainty, and disagreement is more likely to travel to the people who need it — nothing more, and, properly understood, nothing less.

How Leadership Behavior Shapes the Climate

If psychological safety is a property of team climate rather than an individual trait, the natural next question is what actually shapes that climate — and the research consistently points toward leader behavior as one of the most influential and most controllable factors, even though, as established earlier, the causal evidence base for what produces psychological safety remains thinner than the evidence for its consequences.

Nembhard and Edmondson's 2006 study of healthcare improvement teams introduced and examined the specific construct of leader inclusiveness — behaviors through which a leader actively invites and appreciates others' contributions, as distinct from simply not punishing dissent when it occurs. Their findings connected leader inclusiveness to psychological safety and, downstream, to staff engagement in quality improvement work (Nembhard & Edmondson, 2006). The distinction between inclusiveness and mere tolerance is an important one for a leadership audience: not actively punishing someone for speaking up is a low bar, and passing it does not reliably produce a climate where people feel safe doing so. Active invitation — visibly asking for dissenting views, responding to bad news without punishing the messenger, and demonstrably changing course based on what's raised — appears to matter considerably more than simple non-punishment. A separate, larger-sample study by Detert and Burris (2007), surveying 3,149 employees and 223 managers in a restaurant chain, tested this more directly: it found that managerial openness specifically — not general transformational leadership style — was the more consistent predictor of whether subordinates raised improvement-oriented concerns, and that this relationship operated through subordinates' perceptions of psychological safety, statistically consistent with safety functioning as the mechanism connecting leader behavior to voice rather than leader behavior affecting voice by some other route (Detert & Burris, 2007).

This connects to a structural pattern worth naming directly, because it is one a leadership team is unlikely to discover on its own: leaders are structurally poorly positioned to accurately judge the psychological safety of the teams below them, for the simple reason that dissent directed at leadership is exactly the category of information least likely to reach leadership if safety is genuinely low. A leader operating in a low-safety environment does not typically receive direct evidence of that fact — they receive an absence of the evidence that would tell them otherwise. This is inference from the structural information asymmetry described above, not a claim this article can point to a single specific, fully verified study to establish — and it is stated as such rather than dressed up with a citation this article cannot stand behind. What the structural logic implies, though, is testable in principle: if leaders systematically receive a filtered, more favorable sample of climate-relevant information than their teams actually experience, a leader's own assessment of psychological safety should be expected to run ahead of the team's, not because of any failure of judgment, but because the input itself is skewed before judgment is ever applied to it.

None of this is a claim that leaders are typically acting in bad faith, or that low psychological safety reflects poor character. The research is better read as describing a structural information problem: the behaviors that create or damage psychological safety are often subtle, cumulative, and only fully visible from the vantage point of the person taking the interpersonal risk — not from the vantage point of the person whose reaction determines whether that risk paid off. This is precisely why an internal, systematic, and honest read on the actual state of the organization — rather than a leader's own impression of it — has genuine value that direct observation alone cannot substitute for.

The Gap Between Belief, Experience, and Disclosure

This is the natural bridge between the research and the practical question of organizational measurement, and it is worth being precise about the layers involved, because collapsing them is where a lot of well-intentioned measurement effort goes wrong.

There are, functionally, four distinct things in play, and they are not the same: what leadership believes about the organization's climate; what employees actually experience day to day; what employees would be willing to disclose about that experience if asked directly, informally, by someone in their reporting line; and what an organization's formal measurement systems actually capture. Each layer can diverge from the one before it, and each divergence has a specific, identifiable cause.

Leadership belief diverges from employee experience partly because of the structural filtering problem described in the previous section — leaders receive a systematically filtered sample of the actual climate, biased toward the more favorable end, because the people most affected by low safety are the ones least likely to volunteer that fact directly to the leader in question.

Employee experience diverges from what employees are willing to disclose for the same underlying reason, applied at the individual level: how someone actually feels about the climate and what they are willing to say about it to a person with power over their role, compensation, or reputation are not automatically the same thing, particularly for exactly the categories of information — mistakes, disagreement, uncertainty — that this entire body of research concerns itself with.

And what employees are willing to disclose informally diverges again from what a formal measurement system captures, depending entirely on how that system is designed. A survey administered by, visible to, or traceable back to the direct management chain reproduces the same disclosure problem it is meant to solve, because it is asking people to be honest about their willingness to be honest, using the exact channel their honesty is supposed to be about.

This is the precise reason a structurally anonymous, aggregated measurement approach has genuine methodological value here, distinct from any product positioning: it is specifically designed to interrupt this chain of degradation at the third link — between actual experience and willingness to disclose — by removing the interpersonal cost that ordinarily suppresses disclosure. It is worth being equally clear about what this kind of measurement does not do. It does not directly observe the first layer (actual day-to-day experience) — it observes self-reported perception of that experience, filtered through whatever a specific set of survey questions asks. And it reveals a signal, aggregated across many people, about the current state of a specific, defined construct — not the complete organizational truth, and not a diagnosis of why the signal looks the way it does.

What Organizational Measurement Should Actually Target

Given the definitional precision the earlier sections established, a useful measurement approach should target specific, observable behavioral indicators connected to the construct as the research actually defines it — not a generic "how happy are you" sentiment check, which measures something closer to comfort or satisfaction than to psychological safety specifically.

The research points toward several more precise targets. Whether people raise a concern or a near-miss before being specifically asked about it — a direct behavioral indicator connected to the error-detection mechanism at the center of Edmondson's original work. Whether disagreeing with a decision carries a lasting social cost, which targets the interpersonal-risk core of the construct directly, distinct from a general question about whether disagreement happens at all. Whether junior or newer team members challenge senior colleagues' thinking as readily as senior colleagues challenge each other — a specific test of whether safety is evenly distributed across the hierarchy or concentrated among people who already have positional security, which connects directly to the leader-inclusiveness findings discussed earlier. And whether admitting uncertainty or confusion in front of others carries a visible cost, which targets the learning-behavior mechanism specifically, since asking a clarifying question is one of the most basic learning behaviors the original research examined.

Each of these is a meaningfully different, more falsifiable question than a general "do you feel safe here" item, and each is chosen because it maps to a specific mechanism the peer-reviewed literature discussed in this article has actually examined — not because it sounds intuitively related to the topic.

Where This Research Meets Organizational Measurement in Practice

The preceding sections describe a real methodological problem: psychological safety is a genuine, well-researched organizational condition; leaders are structurally poorly positioned to assess it directly; and the layers between actual experience and formal measurement each introduce their own distortion. Klarwerk Insights' Company Radar assessment was built specifically around addressing the third and most tractable of those layers — the gap between what people experience and what they're willing to disclose through a formal channel.

Trust & Psychological Safety is one of Company Radar's eight measured dimensions, and it is the single most heavily weighted input to the platform's Risk Index, contributing 45% of that index's calculation — reflecting the same logic this article has developed throughout: whether an organization can detect a developing problem early depends heavily on whether the underlying climate allows that problem to be raised in the first place.

The specific survey items in this dimension were designed to target the precise behavioral indicators discussed in the previous section, rather than a generic comfort or sentiment measure: whether people raise a mistake or a near-miss before being asked about it; whether disagreeing with a decision carries a lasting social cost; whether newer or more junior employees challenge senior colleagues' thinking as readily as senior colleagues challenge one another; and whether admitting confusion in a meeting affects how someone is perceived afterward. Each of these maps directly to a mechanism discussed in this article — error detection, interpersonal risk, hierarchical distribution of safety, and learning behavior, respectively.

Because Company Radar aggregates responses across an entire organization and only reports results for groups of five or more respondents, with no technical link between a participant's identity and their individual answers, it is specifically structured to interrupt the disclosure-suppression problem described earlier — without claiming to eliminate every form of it. A department comparison can reveal that psychological safety, as measured by these specific items, is markedly lower in one function than in the rest of the organization, converting an abstract company-wide concern into something locatable and specific enough to investigate directly. Where the underlying data supports it, the platform's Executive Tensions logic can also surface a specific, disclosed pattern — for instance, a case where an organization's Execution Index is materially stronger than its Trust & Psychological Safety score, consistent with an organization that may be executing quickly in a way that is outpacing its own ability to detect a developing problem early. This is presented explicitly as a pattern worth an executive conversation, not as proof of a specific cause.

It is equally important to state plainly what this measurement cannot do, consistent with the honesty this entire article has tried to model. It cannot determine why a specific department's psychological safety score is lower than the rest of the organization — that requires a human conversation, not a survey result. It cannot identify which individual is responsible for a low score, and the platform's anonymity architecture makes that technically impossible by design, not merely a matter of policy. It cannot distinguish a temporary dip tied to a single difficult quarter from a persistent structural condition without repeated measurement over time. And it cannot establish that a low score is causing any specific downstream business outcome — consistent with the causal limitations in the underlying academic literature this article has been explicit about throughout. What it can do is convert a condition that is, by its nature, difficult for leadership to observe directly into a disclosed, aggregated, repeatable signal — which is a meaningfully different and more modest claim than "revealing the truth about your organization," and a more honest one.

A Hypothetical Illustration

To make this concrete without overstating what any single assessment can prove, consider a hypothetical scenario modeled on the kind of pattern this article has discussed throughout — not a real customer, but an illustrative composite consistent with how the underlying mechanisms interact.

A logistics company completes its first Company Radar assessment. Its Execution Index comes back high — strategic priorities are being translated into daily work, and cross-team processes are functioning well in practice, not just on paper. Its Risk Index comes back meaningfully lower, driven primarily by a below-average Trust & Psychological Safety score, with Communication also somewhat below the organization's other dimensions. Company-wide, no other single metric looks alarming.

Read in isolation, the Trust & Psychological Safety score is a signal, not a story. Read alongside the Execution Index, per the tension logic described in the previous section, it becomes a more specific and more useful pattern: an organization that is executing quickly in a way that may be outpacing its own capacity to notice a developing problem before it becomes expensive. This is precisely the kind of combination a leader reviewing dimension scores individually, without the cross-dimension view, would be unlikely to notice on their own.

The responsible next step is not to conclude that a specific manager is suppressing dissent, or that a specific policy is the cause — the assessment does not, and structurally cannot, support either of those conclusions. The responsible next step is a leadership conversation grounded in a specific, falsifiable question: if something were currently going wrong inside this organization, would leadership find out about it in time to act, or only after it had already become costly? That is a question a CEO can take directly into a room with two or three people several levels below leadership, framed as genuine curiosity rather than an audit — and the research reviewed in this article, particularly Edmondson's original finding that a unit's reported error count is not a reliable stand-in for its actual rate of mistakes, is exactly why that conversation is worth having even when every visible metric currently looks fine.

What the CEO Should Ask Next

The following questions are intended to be brought directly into a leadership conversation, not answered from a desk. Each is built to be more specific and more falsifiable than a generic "how is our culture" question, consistent with the behavioral, mechanism-specific framing this article has used throughout.

1. When was the last time someone on this team told you something you didn't want to hear — and what happened immediately afterward, in the room?

2. If a serious mistake happened on your team next week, how confident are you that you would find out about it within a day, rather than within a month?

3. Think of the newest or most junior person on your team. When did they last openly disagree with someone more senior — and if you can't think of an example, what does that tell you?

4. Is there a decision made in the last quarter that, in hindsight, someone probably saw coming but didn't say anything about beforehand?

5. When someone admits in a meeting that they don't understand something, what actually happens to how the room treats them for the rest of that meeting?

6. Where in the organization would a piece of bad news currently take the longest to reach you — and why there specifically?

7. If you asked your team, privately and anonymously, whether raising a concern here carries any risk to how they're seen, are you confident you already know what they'd say?

What This Kind of Assessment Cannot Determine

Intellectual honesty about the limits of both the underlying research and any organizational measurement built on it is not a defensive footnote — it is the difference between a genuinely useful diagnostic and an overclaimed one, and it is worth stating these limits with the same directness as everything else in this article.

The underlying academic research, even at its strongest, establishes consistent correlational patterns across large aggregated samples — not proof that psychological safety causes any specific outcome in any specific organization. The field's own leading meta-analysis is explicit that the evidence base for what causes psychological safety to develop, as distinct from what it correlates with, remains thin due to a scarcity of experimental research (Frazier et al., 2017). Any organizational measurement built on this construct inherits that same limitation and cannot resolve it through better survey design alone.

A Company Radar assessment specifically cannot determine why a given score is what it is. A low Trust & Psychological Safety score in a specific department is consistent with several different underlying situations — a specific manager's behavior, a recent difficult event, a longer-standing structural condition, an unusually demanding period unrelated to any individual's conduct — and the assessment has no mechanism for distinguishing between them. That determination requires a human conversation, conducted with judgment and context the survey does not have access to.

It cannot identify individual responsibility, and this is a structural fact rather than a design preference: the platform's anonymity architecture means no individual-level response exists anywhere in the underlying data for any score to be traced back to. This is true even in principle, not merely as a matter of access control.

It cannot fully account for hidden individual circumstances that might affect a particular result — a department going through an unusually difficult period for reasons entirely outside its control, or a small group whose aggregate score is disproportionately shaped by a handful of people's temporary circumstances rather than a stable organizational condition. This risk is one of the reasons aggregation thresholds and repeated measurement over time matter more than any single number from any single cycle.

And it cannot establish, on the basis of one assessment, whether a given finding represents a genuine structural condition or a temporary fluctuation. This is not a flaw specific to this platform — it is a mathematical limitation of any single measurement, academic or applied, and it is the direct reason the research literature this article has drawn on relies on replication across many studies and samples rather than treating any single finding as conclusive on its own.

None of this undermines the value of measurement — it defines what that value actually is. A well-designed assessment converts an otherwise invisible condition into a disclosed, falsifiable signal worth investigating. It does not, and should not claim to, replace the investigation itself.

From Insight to Action

The engineer at the start of this article did not stay silent because she was disengaged, or because leadership had failed her in any dramatic way. She stayed silent because, in a specific room, at a specific moment, the calculation favored silence — and that calculation is precisely what this entire body of research is about.

This is, in one sense, a difficult thing for a leadership team to sit with: the organization's honest self-knowledge depends on a psychological condition that leaders are structurally poorly positioned to assess from where they sit, and that most organizations have no systematic way of measuring at all. But the more accurate and more useful way to hold that fact is not as a source of alarm. It is a solvable measurement problem, not a permanent blind spot.

Silence is not proof that something is wrong. It is, at most, an absence of information — and the research reviewed here is precise about the difference between the two. Psychological safety is not something a leadership team can simply hope exists, and it is not something that has to remain a guess, either. It is something that can be investigated — not perfectly, not completely, but with real, replicated evidence behind the attempt — because the conditions determining whether people speak or stay silent are identifiable, they vary meaningfully across teams and organizations, and they are directly observable once measured honestly and aggregated at a scale that protects the people being asked to be honest.

That is the entire, modest, and genuinely useful promise here: not a guarantee of knowing everything happening inside an organization, but a disciplined way of finding out more of it, sooner, than leadership would otherwise. And once that signal is visible, leadership has something it did not have before: a specific opportunity to respond, while the underlying condition is still cheap to address rather than after it has already cost something real. The conversation that follows — with real people, using real judgment, informed by a signal rather than replaced by one — is where the actual value gets created. The measurement is only ever the beginning of that conversation, not a substitute for having it.

References

Foundational Academic Research

  • Kahn, W. A. (1990). Psychological Conditions of Personal Engagement and Disengagement at Work. Academy of Management Journal, 33(4), 692–724. DOI →
  • Edmondson, A. (1996). Learning from Mistakes Is Easier Said Than Done: Group and Organizational Influences on the Detection and Correction of Human Error. Journal of Applied Behavioral Science, 32(1), 5–28. DOI →
  • Edmondson, A. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly, 44(2), 350–383. DOI →

Empirical Research

  • Nembhard, I. M., & Edmondson, A. C. (2006). Making It Safe: The Effects of Leader Inclusiveness and Professional Status on Psychological Safety and Improvement Efforts in Health Care Teams. Journal of Organizational Behavior, 27(7), 941–966. DOI →
  • Detert, J. R., & Burris, E. R. (2007). Leadership Behavior and Employee Voice: Is the Door Really Open?. Academy of Management Journal, 50(4), 869–884. DOI →
  • Edmondson, A. C., Bohmer, R. M., & Pisano, G. P. (2001). Disrupted Routines: Team Learning and New Technology Implementation in Hospitals. Administrative Science Quarterly, 46(4), 685–716. DOI →
  • Baer, M., & Frese, M. (2003). Innovation Is Not Enough: Climates for Initiative and Psychological Safety, Process Innovations, and Firm Performance. Journal of Organizational Behavior, 24(1), 45–68. DOI →
  • Kessel, M., Kratzer, J., & Schultz, C. (2012). Psychological Safety, Knowledge Sharing, and Creative Performance in Healthcare Teams. Creativity and Innovation Management, 21(2), 147–157. DOI →
  • Carmeli, A., & Gittell, J. H. (2009). High-Quality Relationships, Psychological Safety, and Learning from Failures in Work Organizations. Journal of Organizational Behavior, 30(6), 709–729. DOI →

Reviews / Meta-Analyses

  • Edmondson, A. C., & Lei, Z. (2014). Psychological Safety: The History, Renaissance, and Future of an Interpersonal Construct. Annual Review of Organizational Psychology and Organizational Behavior, 1(1), 23–43. DOI →
  • Newman, A., Donohue, R., & Eva, N. (2017). Psychological Safety: A Systematic Review of the Literature. Human Resource Management Review, 27(3), 521–535. DOI →
  • Frazier, M. L., Fainshmidt, S., Klinger, R. L., Pezeshkan, A., & Vracheva, V. (2017). Psychological Safety: A Meta-Analytic Review and Extension. Personnel Psychology, 70(1), 113–165. DOI →

Institutional / Applied Research

  • Google re:Work (2015). Guide: Understand Team Effectiveness (Project Aristotle). Google re:Work. Source →Internal, observational corporate research — not peer-reviewed. Cited as an applied case only; see Section D for methodological caveats.
  • Duhigg, C. (2016). What Google Learned From Its Quest to Build the Perfect Team. The New York Times Magazine, February 25, 2016. Source →Journalistic account of Project Aristotle, not an academic source.

Ready for your first Company Radar

Organizational Intelligence emerges when every perspective matters.

Bring your organization in, invite your team, and receive strategic intelligence leadership can act on — fully anonymous, within 48 hours.