If you maintain an MCP server, at some point you ask this. You've got twenty tools, you're about to add five more, and something feels wrong about it — but you can't say what, and there's no guidance anywhere.
I read the tool schemas of 4,951 public MCP servers to answer it. 87,146 tools, 270,487 parameters.
The short answer: around thirty. Past that, the thing that breaks isn't your server. It's whether a model can tell your tools apart.
The measurement
For every tool, I took the content words in its description and counted how many appear in no other tool's description on the same server. Call it the tool's distinctive share.
If that number is zero, every word in the description is a word its siblings also use. A model choosing between your tools has nothing in the descriptions to choose on — the names are doing all the work.
Then I split the corpus by how many tools each server publishes.
| Tools on the server | Servers | Zero-distinctive tools | Parameters with no description |
|---|---|---|---|
| 1–3 | 1,256 | 0.5% | 14.8% |
| 4–7 | 1,292 | 1.6% | 22.4% |
| 8–15 | 1,044 | 4.3% | 21.9% |
| 16–30 | 767 | 7.7% | 24.8% |
| 31–60 | 377 | 16.3% | 22.2% |
| 61+ | 215 | 31.3% | 20.5% |
A factor of sixty, rising monotonically. On servers with more than sixty tools, nearly one tool in three has no distinguishing word at all.
Why it gets worse
Part of it is arithmetic. More tools means more chances that two of them collide, and that would happen even if every author wrote carefully.
But the curve steepens around thirty, and arithmetic alone doesn't explain that. What happens around thirty is that authors stop writing descriptions one at a time and start generating them from a pattern. A template is a machine for producing tools that read alike.
The most extreme case in the corpus: one server appends the same 51-word context block to all 275 of its tools.
3land_createCollection "Create a new NFT collection on 3.Land marketplace.
SAP MCP context: Protocol 3land; operation class
write. Use for 3.Land NFT collection, minting,
listing, cancellation, and purchase flows…"
3land_buyNFT "Purchase an NFT from a 3.Land listing.
SAP MCP context: Protocol 3land; operation class
write. Use for 3.Land NFT collection, minting,
listing, cancellation, and purchase flows…"
Creating a collection and buying one are different operations, and the opening sentence says so — in 8 words out of 59. The other 51 are identical across both, and across all 275.
That block was added deliberately, to help.
Zero doesn't mean badly written
This is the part worth internalising, because it's counterintuitive.
Four tools from a widely-installed Gmail server:
Gmail_DeleteDraftEmail "Delete a draft email using the Gmail API."
Gmail_SendDraftEmail "Send a draft email using the Gmail API."
Gmail_ListLabels "List all the labels in the user's mailbox."
Gmail_SearchThreads "Search for threads in the user's mailbox."
Every one of those is clear, correct English. No reviewer would flag them. Every one is also built entirely out of words the other tools use — delete, draft, email, gmail, api, list, search, threads, mailbox all recur across the set.
The description tells you what the tool does. It doesn't tell you what this tool does and the others don't. That second thing is what a model needs at the moment it's choosing, and it's a different question from "is this description good."
The column that argues against splitting
Look at the right-hand column again. Parameters with no description at all sit between 20% and 25% at every size above the smallest bucket. It doesn't improve as servers get smaller.
A two-tool server has the habit about as much as a two-hundred-tool one.
So the two failures are independent. Splitting a large server reduces your description collisions and does nothing whatsoever for your undescribed parameters. They need separate fixes, and the split only buys you one of them.
So should you split?
Split when the collisions are real, not because you crossed a number.
The threshold in the data is around thirty, but that's a population average and your server isn't the population. A server with forty tools that all do genuinely different things to genuinely different objects may be fine. A server with twelve tools where four of them are variations on "search" is not.
The test that actually tells you: read your tool list as one block, the way a model receives it. Nothing else. No README, no repo, no memory of what you meant. Then ask which tool you'd pick for a request that could plausibly go to two of them.
Three options when the answer is "I can't tell":
Rewrite for contrast rather than clarity. Not "is this description clear" but "is it clear which of my tools this is." If you have a shared preamble on every tool, it's costing more than it's buying — the distinguishing sentence shouldn't be a seventh of the text.
Collapse near-identical tools into one with a mode parameter. Four search variants become one search with a scope enum. Fewer things to choose between, and the choice the model has to make moves from "which tool" to "which value," which an enum can constrain and a description can't.
Split the server. Fewer tools per connection, more servers to maintain. Worth it when the tools genuinely belong to different domains, less so when you're just cutting an arbitrary list in half.
The limits of this
This is static analysis. I never ran a model against any of these servers, so I can't tell you how often collisions actually cost anything. A zero-distinctive description might be harmless when the tool name is unambiguous, and expensive when it isn't. Ranking these signals by how well they predict a real mistake needs a model in the loop, which is the next study.
One more hole, found by a server author after I published: the undescribed-parameter figure measures absence only. A parameter described as "query: The query" counts as described and passes. Restating the parameter name is arguably the more common failure and it passes every linter, so 21.8% is a floor.
Data and analysis scripts: https://github.com/getmcpulse/mcp-schema-study
If you want your own numbers rather than the corpus averages, there's a free checker at https://getmcpulse.com/check — paste your tools/list JSON and it scores your distinctive share, undescribed parameters, and token cost against all 4,951 servers. Browser only, nothing uploaded.
I'm building MCPulse, an SDK that reports what models actually do with your tools once real traffic arrives. Everything above came from outside the server, which is exactly its limit: a schema can tell you a model has nothing to choose on, but only traffic tells you whether it chose wrong.
Originally published at getmcpulse.com.
Top comments (19)
One adjustment to the distinctive share measure that I think would sharpen the curve. On a single-domain server the domain noun appears in every description by design. A memory server says memory in all of its tools, a ticket server says ticket, and those words are counted as shared even though they are not a collision at all. They are the name of the server repeated. If words that appear in, say, more than eighty percent of a server's tools were treated as the domain and dropped before counting, small focused servers would score higher and the drop you see past thirty would probably get steeper, because large servers tend to span several domains at once.
Once the domain noun is gone, what is left to distinguish tools on a focused server is mostly two things: the verb, and the trigger condition. Search versus list the most recent versus report counts. Call before acting versus call after a decision was made. The trigger sentence is also the part most often generated from a template, which fits your observation about what happens around thirty.
A second axis your corpus might already contain: tool annotations. Two tools with near identical descriptions but different readOnlyHint or destructiveHint values are distinguishable to the host, which can auto-approve one and prompt for the other, while staying indistinguishable to the model choosing between them. It would be interesting to see how often the zero-distinctive pairs differ only in their annotations.
The domain-noun correction is right and it's a bias I built in without noticing. A memory server saying "memory" in all twelve descriptions is penalised for being focused — the word is the server's name repeated, not a collision. My metric treats consistency as ambiguity.
The 80% threshold is a clean way to strip it, and your prediction about the curve is testable from data I already have. Two competing effects, and I don't know which wins: small focused servers should score higher once the domain word is removed, while large servers spanning several domains have no single word hitting 80%, so nothing gets stripped and their score barely moves. If that's what happens the gap widens and the thirty-tool inflection gets sharper. If it doesn't, my read of what drives the curve is wrong. Either way it's a better measurement than what I ran.
Verb plus trigger condition as what remains is the sharpest framing in this thread. The verb is usually there. The trigger — call this before acting, call that after a decision is made — almost never is, and you're right that it's the part a template flattens. A generated description will carry the verb because the verb is the tool name. It won't carry the condition, because the condition is the only part that requires thinking about the other tools.
Annotations: yes, and someone raised the same axis on another thread, which makes two independent votes. Zero-distinctive pairs that differ only in readOnlyHint or destructiveHint is a query I can run this week — they're in the registry payload. And the asymmetry you name is the interesting part: the host can auto-approve one and prompt for the other, so it can tell them apart, while the model choosing between them cannot. A safety distinction the client honours and the chooser can't see.
Careful, because this adjustment and the "input overlap" one fight each other, and our server is the case where you can see it.
Ten tools, one domain. The nouns that appear in most descriptions are
resume(4) andjob(3). Drop anything over an 80% threshold and you lose nothing here, but lower it toward "the domain noun" and you delete exactly the words that makescore_resume,optimize_resumeandanalyze_job_descriptionconfusable. The collision on a focused server IS the domain noun, repeated on tools that then fail to say how their outputs differ.So I would keep both, as separate readings rather than one corrected score:
A focused server should score high on the first and can still be a mess on the second. That gap is the thing worth reporting, and averaging them into one number hides it.
Your second half is the part I would build on though. The trigger condition, call before or after acting, is the only thing that separates our read-only search from the write that replaces the same list, and no lexical measure sees it at all. We ended up writing "Read-only" and "REPLACES" in capitals by hand because of that.
This is the thing I'd have got wrong, and it would have been silent. Strip the domain noun globally and the input-overlap signal loses precisely the words it runs on — the collision on a focused server IS the domain noun repeated across tools that then fail to say how their outputs differ. I'd have shipped a "sharper" metric that quietly broke the more useful one.
Two readings rather than one corrected score, with different preprocessing feeding each. Domain-stripped for "do these look alike", domain-intact for "could both be right for the same sentence". And the gap between them is the report, not an artefact to reconcile: high on the first and bad on the second is a focused server whose tools all act on the same object and never say what they return. Which is your server, and it's a common enough shape that it deserves its own line rather than an average.
That also resolves something from another thread — someone kept a pair with identical input shapes and clearly different outputs, and was right to keep it. Input overlap alone would flag them. Input overlap plus output indistinguishability is the actual failure, and the two-reading split is where that lives.
On the trigger condition: no lexical measure sees it, agreed, but annotations are a partial proxy. Your read-only search and your replacing write differ on readOnlyHint and destructiveHint whether or not the prose says so. So "pairs that are lexically indistinguishable but differ in annotations" finds the specific case where the host can tell them apart and the model can't — which is your capitals, except machine-readable and already in the payload.
It doesn't catch "call this before acting, that after a decision", which has no annotation and probably can't have one. That part may just be a writing rule rather than a measurement.
Agreed, and your numbers show where the line sits. On your server resume is in four of ten descriptions and job in three, so an eighty percent threshold never touches them, while on a memory server the word memory is in every tool. The threshold is less a knob to lower than a boundary between two things the same count already sees: a noun in nearly every tool is the name of the server, and a noun in a subset of the tools is a cluster, with the subset as the list of candidate collisions.
Run that on your ten and the tightest group is ats, shared by score and optimize and by nothing else, which is the pair you flagged first. Analyze is the instructive case. In the descriptions its only link to optimize is screens against screening, which a counter that does not stem will miss, and its link to score is nothing at all. It joins them through the request, check my resume for this job, which names both objects at once. That looks like the limit of any measure computed on descriptions alone: the collision lives between a sentence and the tools, so the last reading needs a handful of real requests and not only the schema.
The reframe is better than the threshold. Frequency isn't a dial to tune, it's a distinction the same count already makes: a noun in nearly every tool names the server, a noun in a subset names a cluster — and the subset is the candidate list, which is exactly what I was trying to compute separately.
The ats observation is the sharpest test of it. Shared by score and optimize and nothing else, which is the pair flagged first. Two tools out of ten is a cluster; ten out of ten would be the domain. Same counter, no threshold.
And the stemming point is a genuine bug rather than a nuance. screens against screening fails on exact match, so my counter scores analyze and optimize as unrelated when they share their most meaningful word. That's not the interesting finding here — it's just wrong, and it's wrong everywhere in the corpus. Porter stemming before counting is a one-line fix and I don't know yet how much it moves the headline number.
The analyze-to-score case is the one I can't reach. No shared word, stemmed or not, and they collide anyway — because "check my resume for this job" names both objects and the tools partition an intent that the sentence doesn't. The collision exists between a request and a set of tools, not between two descriptions, so nothing computed on descriptions alone will ever see it. That's a hard ceiling on the static approach and I should say so in the post rather than treating it as a gap to close.
Which makes the last reading a small labelled set of real requests. Not traffic at scale — a handful of plausible sentences and which tool each should route to. That's the only artefact that catches the analyze case, and it's the one thing I can't generate from a corpus.
Saying the ceiling out loud in the post is the right call, and it changes what the static metric is for. If the analyze-to-score collision lives between a request and a set of tools, then no amount of work on descriptions will find it, but the description metric is still the cheapest way to decide which pairs are worth testing. It stops being the answer and becomes the filter that produces the candidates.
The labelled set does not have to come from traffic, and it is smaller than it sounds. For each candidate pair, write two sentences: one that should route to exactly one tool, and one that could fairly route to either. The first is the control. The second is the probe. Run each a handful of times at fixed temperature. If the model sends the probe sentence to the same tool every time, the descriptions are distinguishable to it even where the words overlap. If it splits, the pair is a real collision, and you have the sentence that proves it.
That also gives the analyze case a shape you can report. "Check my resume for this job" is not a failure of either description, it is a sentence that names both objects the tools partition. The honest fix there is usually not better wording but a single tool with a mode, which is the line from the top of this thread, arrived at from the other direction.
The control-and-probe pair is the design detail I was missing, and it's what makes the labelled set cheap enough to actually build. I'd been picturing a broad intent corpus, which is a research project. Two sentences per candidate pair is an afternoon, and the control does most of the work — if the unambiguous sentence doesn't route correctly, the problem isn't ambiguity between the pair, it's that one description is broken on its own terms, and that's a different fix.
A stable split is the thing I'd have misread. My instinct would have been to treat any non-uniform distribution as a collision, but consistent routing under an ambiguous probe means the model has found a distinction I can't see in the words — which is a pass, not a failure. Only the split is evidence. Fixed temperature and a handful of runs is enough to tell those apart, which is a much smaller budget than I'd assumed.
And it reframes the metric usefully. Static analysis stops being the answer and becomes the thing that produces candidates worth spending model calls on. That's the version that survives at scale — you can't probe 191 tools pairwise, but you can rank pairs by input overlap and probe the top fifty.
The analyze case landing on "one tool with a mode" is the part I like most, because it arrived from the opposite end of the thread. The rule came from the interface — if two tools can ever be correct for the same sentence, you have one tool with a parameter. The probe gets there empirically, and hands you the sentence that proves it. Same conclusion, and the second version is the one you can put in a pull request without arguing.
The distinguishability framing is the useful part, and it matches what we hit from the other side.
We run a remote MCP server with ten tools for job applications, and the count was never the constraint. The thing that decided whether a model picked the right one was how much the descriptions overlapped in their verbs. "score a resume", "analyse a job description" and "prepare for an interview" all read as "look at this posting and tell me something", and the model would happily reach for any of the three. Rewriting the descriptions so each one names its own output rather than its own input fixed more than trimming the list would have.
One thing I would add to your thirty: it is not just the count, it is whether two tools can ever be correct for the same sentence. If they can, you do not have two tools, you have one tool with a parameter.
Ours is at aiapplyd.com/mcps if you want a ten-tool sample for the dataset.
"If two tools can ever be correct for the same sentence, you have one tool with a parameter" is a better rule than the number I led with. It's a property of the interface rather than a property of the list, and you can check it without any data.
The verb point is the sharper half of it though, and my metric partly misses it. I counted content words, so "score", "analyse" and "prepare" are three distinct tokens and all three tools score as distinctive. Your set would look fine by my measurement and still be ambiguous in practice, because the verbs are different words for the same act — read this posting, tell me something about it. Lexical distinctness isn't semantic distinctness, and I only measured the first.
Naming the output rather than the input is the fix and I want to steal the phrasing. Every tool that takes a job posting has the same input, so describing the input is describing what they share. The output is the only place they differ, which makes it the only place a model can discriminate. That generalises past your domain — any set of tools operating on the same object has this shape, and it's exactly where my corpus showed the worst collisions.
Which suggests a measurement I should have run: distinctive share computed on verbs and output nouns separately, rather than over all content words at once. A tool can be 40% distinctive overall and 0% distinctive on the part that decides selection.
Yes to the sample, thank you — ten tools on one object is a better test case for that than anything in the corpus.
Here it is. Ten tools, all working on one object: a job posting, or the resume held up against it. First sentence of each description, as the model sees it:
aiapplyd_score_resume: Score a resume for ATS compatibility.aiapplyd_analyze_job_description: Extract what a job posting actually screens on.aiapplyd_optimize_resume: Rewrite a resume so it passes ATS screening.aiapplyd_generate_interview_questions: Produce interview preparation for a specific role and company.aiapplyd_generate_cover_letter: Write a cover letter for a specific job.aiapplyd_translate_resume: Translate the saved resume into another language.aiapplyd_search_jobs: Return the user's job matches. Read-only.aiapplyd_update_job_preferences: Set the roles and locations it hunts for. Replaces, never appends.aiapplyd_auto_apply: Apply to one specific posting end to end, on the employer's own hiring system.aiapplyd_build_pdf: Build a formatted resume in the builder.Where I'd expect your verb and output split to light up: score, analyze and optimize are the tight cluster. Same input, and "check my resume for this job" is a fair sentence for all three. Only the outputs separate them: a number, a keyword list, a rewritten document. The other pair to watch is search vs update_job_preferences, a read and a write on the same noun, which is why those two descriptions carry "Read-only" and "REPLACES" in capitals.
The server is public at mcp.aiapplyd.com/mcp if you'd rather pull the live descriptions than trust my paste, and there's an overview at aiapplyd.com/mcps?ref=devto. Curious what the verb-only share comes out at.
Ran it. Your set breaks my proposed fix, which is more useful than if it had worked.
Overall distinctive share, lowest first:
50.0% score_resume score, compatibility
60.0% analyze_job_description extract, actually, screens
60.0% optimize_resume rewrite, passes, screening
60.0% generate_cover_letter write, cover, letter
75.0% auto_apply / build_pdf / search_jobs
80.0% translate_resume
83.3% generate_interview_questions
100.0% update_job_preferences
Nothing scores badly. Your tight cluster — score, analyze, optimize — sits at 50–60%, which my corpus would call healthy.
And the verb split doesn't help: all ten verbs are distinct. score, extract, rewrite, produce, write, translate, return, set, apply, build. 10/10 unique. So computing distinctiveness on verbs alone gives every tool a perfect score, including the three you flagged.
Which kills the measurement I described in my last comment. Worth saying plainly.
What does show the cluster is the shared input nouns:
resume in 4 descriptions
job in 3
specific in 3
ats in 2
posting in 2
That's the signal. Not low distinctiveness — high input overlap. Your three cluster tools are the ones whose descriptions share the object being acted on, and they're ambiguous despite being lexically distinctive, because "check my resume for this job" is a fair sentence for all three and nothing in the wording resolves which.
So the metric I should add isn't a finer-grained distinctive share. It's a second, separate signal: which tools share their input noun. Distinctive share tells you whether two descriptions look alike. Input overlap tells you whether two tools could be correct for the same request — which is the thing that actually matters, and it's the measurable version of the rule you gave me.
On your defensive capitals: "Read-only" and "REPLACES" are doing work my metric can't see at all, since they're constraints rather than content. That you had to add them by hand is itself the finding.
Thanks for the list. It changed the fix.
"Input overlap, not low distinctiveness" is the version I would keep. It also explains why the fix that worked on our side was renaming toward the output rather than trimming words: it changes the object in the sentence, and the object is exactly what your second signal measures.
Two things I would add from this end.
The overlap edge probably wants a direction. Ambiguity is rarely symmetric. "Check my resume for this job" should not land on score and optimize equally often, and whichever one wins is the one whose description is doing the persuading. If your corpus can see which tool actually got called on an ambiguous request, the metric stops saying "this pair is confusable" and starts saying which of the two descriptions to rewrite. That is the difference between a diagnosis and a task.
The capitals point at a gap with a real home in the spec. MCP has tool annotations for precisely that class:
readOnlyHint,destructiveHint,idempotentHint,openWorldHint. We set them and still write "Read-only" and "REPLACES" in the prose, because the model reads the description while the client decides what to do with the hints, and those two audiences are not the same. If you can read annotations, a third row of "tools whose constraints exist only in prose" would find a lot of servers, ours included.Thanks for running it. A negative result you publish is worth more than the fix that survives quietly.
Directionality is the right correction and it's the difference between a finding and an instruction. A symmetric metric says "these two are confusable" and leaves you with two descriptions and no idea which to touch. Directed, it says which one is winning, and the winner is the one whose wording is doing the persuading — so the loser is what you rewrite, or the winner is what you narrow.
I can't see it from the corpus. The schema survey never observed a call. But the SDK is inside the server, so the traffic side has it: given an ambiguous pair identified statically, which one actually gets called, and at what ratio. A 50/50 split and a 90/10 split are different problems — the first is a coin flip, the second is one description eating the other's requests. I'd been treating overlap as the whole answer and it's only half.
Annotations I hadn't considered at all, and the audience split you describe is exactly why. readOnlyHint and destructiveHint are read by the client deciding whether to prompt for confirmation. The description is read by the model deciding whether to call the thing. Two readers, two channels, and nothing in the spec says they have to agree — so authors like you correctly write it twice.
Which makes "constraints present in prose but absent from annotations" a real check, and a straightforwardly mechanical one: grep descriptions for read-only, replaces, overwrites, deletes, permanent, cannot be undone, then see whether the matching hint is set. Both directions are findings. Prose without annotation means the client can't protect the user. Annotation without prose means the model doesn't know a tool is destructive when it's choosing.
Adding it. Annotations are in the registry payload, so I can run it across the corpus this week rather than waiting on traffic.
This matches what we found while building a small, read-only MCP server for a ticket system.
We deliberately stopped at six tools—not because six is a magic number, but because each tool needs a distinct decision point for the model: search candidates, read one ticket, retrieve a flat batch, load a bounded analysis bundle, or inspect comments and links.
The interesting design review was not “are the descriptions clear?” but “could two tools both be reasonable for the same user request?” Our closest pair is batch retrieval versus an analysis bundle. They share the same input shape, but differ in intended output: flat details for more tickets versus bounded evidence including optional comments, links, and source anchors.
So I would treat your “around thirty” less as a limit and more as an excellent audit trigger. The decisive test is whether a model can infer the intended output and boundary from the tool name, description, and parameter contract alone—without knowing the repository or the author’s intent.
One additional dimension might be worth measuring: not only lexical overlap in descriptions, but overlap in input objects and user intents. Two descriptions can use distinct words and still compete when both operate on the same object for the same vague request.
Especially useful article—this turns “keep the tool list small” into a testable interface-design question.
Your batch-retrieval versus analysis-bundle pair is the cleanest statement of the case I've seen, because you kept it. Same input shape, different intended output, and you decided it was two tools rather than one. Most write-ups of this either merge the pair or never notice it.
What makes it survivable, I think, is that the boundary is a property of the output and you can say it in a sentence. "Flat details for more tickets" versus "bounded evidence with optional comments and source anchors" — a model reading those has something to discriminate on even though the inputs are identical. That's the case my metric gets wrong in the other direction: it would flag the pair on input overlap, and you'd be right to dismiss it.
Which suggests input overlap needs a partner rather than standing alone. Overlapping inputs plus distinguishable outputs is a fine design. Overlapping inputs plus indistinguishable outputs is the failure. Flagging the first is exactly the false positive that gets a signal switched off.
Your decisive test is the better formulation of what I was reaching for: can a model infer the intended output and boundary from name, description and parameter contract alone, with no knowledge of the repo or the author's intent. That's checkable by a person in about a minute per pair, and it's the thing the corpus can't do for you because I can't read intent either.
Six tools with six distinct decision points is a different discipline from six tools that happen to exist. The count is downstream of that, which is why thirty is an audit trigger rather than a limit — agreed, and I'd say it more strongly than the post does.
"The names are doing all the work" is a great way to put it. I noticed the same failure from the caller side with a small set of agent-facing endpoints: once tools overlap, the model stops reading descriptions and pattern-matches on names, and then your error rate is really a naming problem. Your table gives a concrete ceiling to design against. Splitting one fat server into two focused ones is cheaper than writing cleverer descriptions, and your distinctive-share metric explains exactly why.
"The model stops reading descriptions and pattern-matches on names" is the mechanism I was guessing at and couldn't observe. From the schema side I only see that the descriptions have nothing left to discriminate on — not what the model falls back to. Seeing it from the caller side closes that loop.
One complication on splitting. Another commenter here raised the case where it doesn't help: ten tools that all take a job posting, described as "score a resume", "analyse a job description", "prepare for an interview". Different verbs, so they pass my metric, and the model still can't choose — three names for the same act on the same input. Split that server and you've moved the ambiguity, not removed it.
His rule beat my number: if two tools can ever be correct for the same sentence, you don't have two tools, you have one tool with a parameter.
So the table is a warning sign rather than a threshold. Past thirty, collisions are likely enough to audit for. But the audit is whether any two tools could both be right for one request, and that answer decides between splitting, merging and rewriting.
Your naming point also exposes a gap: I never measured names at all. If names are what the model falls back on, name distinctiveness is what I should have measured alongside.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.