Back

Enterpret

Analyze customer feedback at scale.

MCP server URL

https://enterpret.gumstack.com/mcp

Works with

Tools 7

  • Get Organization Details

    Get details about your organization. Fetches organization metadata including name, slug, and dashboard URL. This is useful for understanding your organization context and generating proper links to customer records. Returns: Organization details including name, slug, and dashboard URL

  • Find User Quote

    Find validated verbatim user quotes from Knowledge Graph record IDs. Use this whenever the user asks for quotes, verbatims, snippets, sample customer language, or examples in the customer's words. First call the graph query tool available on the current surface only to fetch candidate NaturalLanguageInteraction record IDs, then pass those IDs here: run_graph_query on v2/admin schema-decomposition surfaces, execute_cypher_query on the legacy auth-scoped surface. Do not use fi.content/nli.content rows as quotes. If this tool returns no validated quotes, do not substitute summaries or paraphrases. Do not treat product, feature, or customer names in quote requests as source filters unless the user explicitly asks for a source or integration. If a narrow source-like filter returns no candidate records or no quotes, broaden the candidate query across sources before concluding no quotes exist. If the user specified a date window, keep record timestamps from the candidate query and state the actual earliest/latest dated records analyzed as the first sentence of the final response before any title, heading, or quote list. Do not merely restate the user's requested window. Preserve numeric citation links only on exact numeric values returned by the graph query tool. Do not reuse a citation URL for computed, summed, rounded, or approximate numbers. Candidate and count rules: - Plan the final answer rows/categories first, then call this tool once per row/category that needs a quote. Do not invent new sections from quote text returned by this tool. - record_ids are the candidate search pool. count is the number of validated quotes to return from that pool, not the number of records to inspect. - Pass 5-10 quote-suitable record_ids when enough candidates exist; use fewer only when the query genuinely returned fewer matching records. - Use count=1 for a broad representative quote request or for each row/category unless the user asks for an exact count or "at least N" quotes. Plural wording such as "quotes" or "examples" does not by itself require count > 1. - For source-specific quote asks, candidate records and final quotes must come from the requested source/source-channel. Criteria rules: - Each criterion targets ONE topic/claim. It may include close synonyms of the same topic, but must not merge independent topics, causes, or outcomes into one long comma-separated string. - Use customer-language evidence terms and close synonyms from the candidate records — not internal taxonomy labels, category names, or the user's literal question phrasing when the customers themselves use different words. - Match the section's intent, not just nouns. Praise quotes need positive sentiment/usefulness/appreciation; complaint quotes need pain or dissatisfaction; comparison quotes must preserve direction. - Preserve explicit action intent. For leave/switch/cancel/stop-spending/churn asks, the criterion must seek quotes where that action intent appears in the customer's own words — a generic complaint, refund frustration, or product bug is not enough. - For source-specific asks (G2-only, Intercom-only, app-review-only, survey-only, Slack-only, etc.), keep the source scope inside the criterion rather than expecting post-hoc filtering to fix a mismatched pool. - For several final rows/categories, use separate parallel calls with ONE single-topic criterion per row/category — never one broad mixed criterion. Good single-topic criteria: ✅ "fees are too high or feel unfair" ✅ "checkout or payment flow is slow or fails" ✅ "AI results feel inaccurate or untrustworthy" ✅ "fast setup is useful or appreciated" ✅ "customer cannot find key UI controls or where to take the action" ✅ "frustrated with OAuth token expiration during SSO login" ✅ "would switch to a competitor because of high fees" Bad combined-topic criteria (merge independent topics, return mixed/off-scope quotes): ❌ "fees too high, unfair charges, competitor cheaper, refund delay" ❌ "checkout too slow, lost auction item, rewards lost" ❌ "AI accuracy, wrong counts, duplicate records, inflated numbers, mislabeling" ❌ "hard to find controls, UI navigation, missing feature, hidden button, onboarding confusion" Bad section-intent / direction / action-intent mismatches: ❌ "integration issues" — no polarity, no evidence terms ❌ "checkout is fast" — as a complaint criterion (opposite direction) ❌ "refund frustration" — for a churn/leave-intent ask (no leave/switch/stop-spend intent) ❌ "account bug" — for a competitor-defection ask (no move-to-competitor intent) Repair and admission rules: - If this tool returns zero or too few quotes for a required row, repair that missing row once by broadening the semantic criteria or candidate records. Keep already-validated quotes; do not restart successful rows. - Final quote quality filtering is omit-only: if a quote is weak, wrong source, wrong speaker, or supports the opposite direction, omit that row or state that no validated quote was found. Do not emit "Quote unavailable" placeholders. Preserve returned quote text and speaker attribution links in final answers. When a social/page/app-review source represents a customer comment on a brand page, do not use the brand/page/account name as the speaker for the final quote. Attribute that quote to the customer/reviewer when available; otherwise use a neutral "Customer" speaker with the returned record link. This is different from official brand-authored posts, which should be omitted when the user asked for customer comments. For support/chat sources such as Intercom or Zendesk conversations, customer quote requests must use the external customer side of the conversation. Exclude Enterpret/support-agent/internal-staff replies even when they explain the product well. If speaker labels identify CUSTOMER versus AGENT/TEAM, choose CUSTOMER. For prose, keep the exact blockquote shape, including the > marker on both the quote line and the attribution line: > "Quote text" > — [Speaker](link) For tables, keep the quote text and speaker attribution link in the same table cell as `"Quote text" — [Speaker](link)`; do not split them into a generic record-link column. If the user asks for a table, include the requested columns such as date, channel/title, and record link, while preserving quote + attribution in one quote cell. The validated attribution link is the quote's navigation. If the user asks for a separate record-link column in a quote table, use the record link already returned in the speaker attribution or the plain record_id; do not replace the speaker attribution link. Args: criteria: What the quotes should be about record_ids: NLI record IDs returned by a previous graph query count: Number of quotes to return, between 1 and 5 user_question: The user's original question for language and speaker intent detection Returns: Validated quotes formatted as blockquotes with speaker attribution and record links.

  • Get Graph Schema

    Get the Knowledge Graph schema structure for a specific organization. Retrieves entity names, relationships, and static properties. IMPORTANT — check this flag in the response: has_source_fields: - If true: call search_graph_fields(query) before each Cypher query to get relevant integration fields (Zendesk, Salesforce, Gong, etc.) - If false: no source-specific metadata, skip search_graph_fields Returns: Schema structure with has_source_fields flag

  • Get Query Examples

    Get sample Cypher queries similar to what you need to build. Returns reference examples showing correct compiler dialect. Some examples include a wrong_cypher field showing a common mistake and why it fails — study these to avoid the same errors. Examples that need a label, relationship or property this tenant's schema lacks are hidden (see hidden_by_schema and warnings). Never reintroduce a label, relationship or field from memory; confirm it in get_graph_schema first. HOW TO USE THE RESULTS: 1. Study the correct_cypher patterns — relationships, COUNT DISTINCT, date functions, MISC filters, category_enum usage 2. If wrong_cypher is present, understand what mistake it shows and avoid it 3. ADAPT the patterns with exact values, entity labels, and field names discovered from search_graph_values, search_graph_fields, and schema results 4. Build your own query — do NOT copy-paste 5. For month/month-range requests with no year, use the most recent fully completed occurrence relative to the current date. If the range is upcoming or in progress this year, use the previous year. 6. For "last month" / "past month" without a named calendar month, use a rolling window: now() - INTERVAL 30 DAY. For daily/weekly/monthly analysis, keep that same filter for every query; group daily/weekly with toStartOfDay/toStartOfWeek, and make monthly the window total unless the user asks for a multi-month trend. 7. For quote requests inside explicit date windows, candidate-record queries should return record_id plus record_timestamp/source/summary. Make the first sentence of the final answer state the actual earliest/latest records analyzed, before any title/heading/quote list; do not only repeat the requested window. 8. For CASE-bucket counts, prefer WITH ... CASE ... END AS bucket, then RETURN bucket, COUNT(DISTINCT ...) so the compiler groups by the bucket. 9. Preserve explicit source filters from the user. If the user says Zendesk/App Store/Gong/G2, do not swap in a different source discovered from exploratory results. 10. "General" and "Generic" taxonomy values are not Miscellaneous. Unless the user explicitly says "miscellaneous", still filter only type != 'MISC' and do not add type != 'GENERIC'. 11. For source metadata prompts, include the source and field intent in search_graph_fields: "appstore rating country", "zendesk tags customer experience", "gong account name call title", "g2 star rating". 12. Resolve taxonomy/source values with search_graph_values before filtering names. If level is unknown, search relevant types and use the returned object_label/property_path; do not assume the level upfront. 13. Prefer taxonomy filters over content keywords when a relevant taxonomy value exists. Use exact equality/IN on canonical names and type != 'MISC' unless miscellaneous is explicitly requested. 14. L1/L2/L3 are product scope. Theme/Subtheme are customer topics/issues. For complaints/topics under product scope, join Theme/Subtheme from the same cft; complaints use t.category_enum CONTAINS 'COMPLAINT'. 15. If no relevant taxonomy value exists, use content fallback: summary/content searches use fi.content; exact/raw wording searches use nli.content. Preserve quoted/exact phrases as one string. 16. For generic "top themes" / "dominant themes" / theme-ranking requests, always run the Theme-level ranking first without Subtheme: join (cft)-[:HAS_THEME]->(t:Theme), filter t.type != 'MISC', and aggregate by t.display_name to get the exact total for each theme. Do not substitute L1/L2/L3 product taxonomy or content keyword buckets unless the user asks for product/feature areas instead of themes. Then, when subthemes are requested, run a second query for the returned theme display names using (cft)-[:HAS_SUBTHEME]->(st:Subtheme), filter st.type != 'MISC', and add only a few top subthemes under each theme in the final response. Do not answer from the first theme query alone when subthemes are requested; run the second subtheme query before producing the final response. Never sum subtheme counts to produce theme totals; use the first query totals. Do not use WITH ... ORDER BY ... LIMIT followed by another MATCH in one query. 17. Theme and Subtheme labels are object-specific: use `display_name` for filtering, grouping, returning, and presentation. L1, L2, and L3 use `name`. Preserve every other object's existing schema-defined field; never switch another object to `display_name` because of this taxonomy rule. 18. For account-health/at-risk/red-account/churn triage based on complaint volume and "top complaint issues/problems/themes" requests, prefer taxonomy-backed examples over fi.content keyword scans: use Theme.category_enum CONTAINS 'COMPLAINT', exclude MISC taxonomy nodes, and keep aggregation at the user's requested level (Theme for themes, Account for accounts). Ask before mixing IMPROVEMENT into a complaint-only request. This rule does not apply to Opportunity/deal risk predictions, which use DealRiskEnvelope. 19. If an example/query must use OR with taxonomy-name predicates, wrap the OR group before adding other filters: `(t.display_name CONTAINS '<term A>' OR t.display_name CONTAINS '<term B>') AND t.type != 'MISC'`. Prefer exact equality/IN from taxonomy search over CONTAINS. 20. For Opportunity/deal-risk requests, structured DealRiskEnvelope predictions are authoritative when a correctly scoped query succeeds, including zero rows. Do not add LLM-inferred risks. Only if that path is unavailable may you return a separate labeled `LLM-inferred potential risks` section; never combine it with predictions. A risk call snippet also requires an active-schema path from the returned DealRiskEnvelope to that call's NLI record. Do not guess the path or substitute generic deal feedback when no linked row exists. WHEN TO CALL THIS TOOL: Call this tool before writing non-trivial Cypher or before run_graph_query when you need a reliable pattern for relationships, filters, grouping, date handling, source metadata, quotes, taxonomy joins, or compiler-supported syntax. Calling it early usually improves query formulation because the returned examples include sample Cypher plus notes that explain the relevant graph shape and common mistakes. Use it especially when: - The user asks for analysis that combines multiple dimensions, such as source + taxonomy + date, sentiment + source, account + theme, or metadata + theme - You are unsure which relationship chain to use, where to place MISC filters, or whether to count feedback records, accounts, users, themes, or opportunities - The query needs tricky syntax such as WITH/CASE buckets, quote candidate retrieval, explicit date windows, OR filters, or source-specific metadata fields - A previous Cypher attempt failed or returned a suspicious result and you need a nearby working pattern before trying again - You have already used search_graph_values, search_graph_fields, or schema context to identify likely fields/values and now need to assemble the Cypher Do not use this tool to discover live values or answer the user's question by itself. The examples are static formulation guidance. Adapt the sample Cypher to the exact field names, entity labels, values, source names, and dates from the current user request and discovery tools. HOW TO WRITE THE INTENT: The `intent` argument is a short natural-language description of the Cypher you want to build. Specificity beats keywords alone. Include the four aspects of your query: - Operation: "count", "top themes", "sample records", "sentiment breakdown" - Filter: "complaints", "from Zendesk", "in last 30 days" - Grouping: "by theme", "by source", "by month" - Special needs: "exclude billing", "with account info" Good intents (specific → better matches): "count complaints by theme from Zendesk in last month" "monthly trend analysis of feedback volume" "sample records with sentiment and source filtering" "top L2 categories with MISC filtering" "exclude specific topics from source analysis" "compare complaint volume before and after a date" Bad intents (too vague → poor matches): "query" / "get data" / "feedback" / "help" Args: intent: Short natural-language description of the Cypher you want to build. Include operation, filters, grouping, and special needs. top_k: Number of examples to return (default: 5) Returns: JSON array of examples with correct_cypher, notes, and optionally wrong_cypher. Current date for relative date resolution: 2026-09-16 UTC.

  • Search Graph Fields

    Resolve source/customer-specific metadata field names for Cypher. Purpose: use this tool when the user refers to a source-specific or customer-specific metadata field and the exact graph property is unclear. It is not a general schema search across all graph fields, and it does not validate field values. Use this as the field-name resolver for source/customer metadata. If the source-field index is unavailable, the tool returns a warning; that warning is not proof that the field or data is absent. Treat it as a tool/index limitation: fall back to the in-context schema/TOON if available, query examples, or a bounded value/schema inspection before saying the field is unavailable. Use for source/customer metadata fields such as: - Zendesk ticket status, tags, queue, or priority. - App Store rating, country, or app version. - Gong call title, account name, or call metadata. - Salesforce churn reason, opportunity fields, or case metadata. - Plan type, bucket, QID/question fields, NPS/CSAT/DSAT score fields, rating fields, or custom segment fields when these are source metadata. Do not use for: - Fixed graph properties such as Theme.display_name, L1.name, NLI.source, language, or sentiment. - General "find any relevant field in schema" discovery. - Value validation. Use search_graph_values for user-provided values. Mandatory pair with value search: - This tool resolves field names only. It does not validate that a user value exists under the field. - If the user asks for a value filter such as "paid teams", "free plan", "spenders", "Qualtrics", "SMB NPS", or "Zendesk", call `search_graph_values` for that value/source as well. - For questions like "Where can I find NPS score data?", call this tool with a focused field query such as "nps score" or "net promoter". Do not answer from schema absence alone until this field probe has been attempted or the active schema already contains the exact score field. If the field search returns a degraded warning, inspect schema/examples next; do not conclude "NPS is not available" from the degraded warning alone. - For mixed asks such as "Qualtrics SMB NPS bucket spenders QID2", split the work: `search_graph_fields` for field labels like "qualtrics bucket" and "qualtrics QID2"; `search_graph_values` for values like "Qualtrics", "SMB NPS"/"NPS", and "spenders"/"spender"/"spending". Do not treat a successful field hit as proof that any of those values exists. HOW TO CALL: Include the source name in your query for better matches. Good: "zendesk ticket status and tags" Good: "zendesk satisfaction rating csat" Good: "zendesk tags customer experience" Good: "gong call title and account name" Good: "appstore rating country" Good: "g2 star rating" Good: "nps score" Good: "qualtrics bucket" Good: "qualtrics qid2" Bad: "billing issues" (too vague, may return wrong source's fields) Bad: "qualtrics bucket survey name SMB NPS spenders" (mixes fields and values) HOW TO USE RETURNED FIELDS: - Returned fields are query candidates for source/customer metadata that may be omitted from the compact graph schema. Do not reject a returned field just because it is absent from get_graph_schema/TOON. - Use exact returned field names in Cypher WHERE and RETURN clauses unless this tool returned a degraded/unavailable warning or the field's entity group is genuinely incompatible with the requested query. - Do not say a returned field is unavailable only because it is absent from compact schema/TOON. - Format: {alias}.{field_name}, where {alias} matches the entity group each field is listed under in the response. NLI/FeedbackRecord fields use nli.*; Account/User fields use the joined alias (see the per-entity prefix block appended to results for the exact MATCH pattern). - Prefer source-specific fields over fi.content search when the field exists. Good (NLI): WHERE nli.zendesksupport_status CONTAINS 'solved' Good (NLI): WHERE nli.zendesksupport_satisfactionrating CONTAINS 'good' Good (Account): MATCH (...)-[:PROVIDED_BY_ACCOUNT]->(acc:Account) WHERE acc.industry CONTAINS 'Retail' Bad: WHERE fi.content CONTAINS 'Acme' (broad, may match irrelevant content) Bad: WHERE nli.industry CONTAINS 'Retail' (industry is on Account, not NLI) - When multiple similar fields exist (e.g., countryiso2 vs country_full_name), prefer the human-readable variant for filtering. - When returned fields directly match the user's requested source-specific dimension, the final Cypher must keep those fields as mandatory predicates or grouping columns. If the exact source-field query returns no rows, report no matching rows for that source-field constraint instead of switching to generic `uf_*` fields, taxonomy/content matches, or broad source-only results as the main answer. RATING/SCORE FIELDS (stored as lists - use CONTAINS, not =): Good: WHERE nli.appstore_rating CONTAINS '1' (filters 1-star reviews) Good: WHERE nli.g2_metadata_starrating CONTAINS '2' (filters 2-star G2 reviews) Good: WHERE nli.zendesksupport_satisfactionrating CONTAINS 'bad' (filters bad CSAT tickets) Bad: WHERE nli.appstore_rating = 1 (fails because ratings are list fields) To combine with other filters, always parenthesize OR groups: Good: WHERE nli.appstore_rating CONTAINS '1' AND (nli.appstore_country_full CONTAINS 'US' OR nli.appstore_country_full_name CONTAINS 'United States') Args: query: Include the source name (zendesk, gong, appstore, etc.) for best results top_k: Number of source groups to return (range: 1-100; default: 15) Returns: Matched source/customer metadata fields grouped by entity and source platform.

  • Run Graph Query

    Execute a Cypher query against Enterpret's Knowledge Graph. Translates your Cypher to SQL and returns structured results. Enterpret's KG unifies 100+ sources (Zendesk, Gong, Slack, AppStore, etc.) into a single graph. Source identity: nli.source. Source-specific metadata fields live on whichever entity the field belongs to — qualify each one with the alias matching its entity group as returned by search_graph_fields (nli.* for NLI/FeedbackRecord, the joined account/user alias for Account/User fields). BEFORE writing Cypher: - For source-specific fields (ratings, tags, account names, plan types, custom metadata), call search_graph_fields only when the exact field is not already available in the active schema. - Call get_query_examples to see proven patterns for similar queries. - Query examples and memory are not value-resolution evidence: before using a literal L1/L2/L3/Theme/Subtheme/Source name in WHERE, resolve it with search_graph_values. If the level is unknown, search relevant types and use the returned object_label/property_path. You MUST follow the Enterpret-specific Cypher rules below. Standard Neo4j Cypher patterns will often fail or return incorrect results due to the KG's structure (relationship fan-out creates duplicate rows). ## LIMIT Only apply LIMIT when the user explicitly requests a subset or examples. No LIMIT — user wants a true total or complete breakdown: - "how many" / "total" / "count" → COUNT query, no LIMIT - "all feedback" / "every" / "locate all" → exhaustive retrieval, no LIMIT - "breakdown by X" / GROUP BY → all groups needed, no LIMIT Good: RETURN COUNT(DISTINCT nli.record_id) AS total Good: RETURN language, COUNT(DISTINCT nli.record_id) AS total ORDER BY total DESC Bad: RETURN COUNT(DISTINCT nli.record_id) AS total LIMIT 100 -- caps the count! Bad: RETURN ... LIMIT 50 -- when user asked for "all", hides results Use LIMIT — user wants a ranked subset or examples: - "top 5" / "biggest" / "most common" → LIMIT N (default 50) - "show me some examples" / "sample feedback" → platform-aware limit: Gong/Zoom calls ~3-5, Zendesk tickets ~10, Surveys ~20-50 ## Counting Always use COUNT(DISTINCT entity.record_id). Never COUNT(*) or COUNT(entity). Graph traversals create duplicate rows: 1 NLI → many FI → many CFT. Good: COUNT(DISTINCT nli.record_id) AS total Bad: COUNT(*) AS total -- inflated by fan-out duplicates Feedback count → COUNT(DISTINCT nli.record_id) User/customer count → COUNT(DISTINCT <user_alias>.record_id) via the user-join path your org's schema exposes (PROVIDED_BY_USER → User or HAS_USER → DerivedUser; see get_graph_schema and search_graph_fields). ## DISTINCT Always use DISTINCT when returning NLI rows for samples or examples. Without it, the same record_id may appear multiple times due to fan-out. Good: RETURN DISTINCT nli.record_id AS record_id, fi.content AS summary Bad: RETURN nli.record_id, fi.content -- duplicates likely ## MISC Filter Every query using Theme, Subtheme, L1, L2, or L3 MUST include: entity.type != 'MISC' This applies to all query types — counts, samples, filtered, exploratory. Good: WHERE t.category_enum CONTAINS 'COMPLAINT' AND t.type != 'MISC' Bad: WHERE t.category_enum CONTAINS 'COMPLAINT' -- includes junk MISC entries For OPTIONAL MATCH: WHERE (t IS NULL OR t.type != 'MISC') Only omit if user explicitly says "include miscellaneous". ## Dates Relative dates — always use INTERVAL, never hardcode: Good: WHERE nli.record_timestamp >= now() - INTERVAL 30 DAY Good expected-close filter: WHERE opp.close_date >= today() AND opp.close_date < today() + INTERVAL 30 DAY Bad expected-close filter: WHERE opp.close_date >= toDateTime('2024-06-01') -- breaks over time Literal dates: toDateTime('YYYY-MM-DD HH:MM:SS') Month or month range without year: use the most recent fully completed occurrence relative to the current date. If the month/range is upcoming or still in progress this year, use the previous year. For "last month" or "past month" without a named calendar month, use a rolling window: now() - INTERVAL 30 DAY. For daily/weekly/monthly analysis over a relative window, keep the same rolling filter in every query; group daily with toStartOfDay, weekly with toStartOfWeek, and make monthly the total for that same rolling window unless the user asks for a multi-month trend. For explicit time windows, preserve or query actual coverage with returned timestamps or MIN/MAX timestamp fields, and state the actual analyzed min/max as the first sentence of the response before any title, heading, or quote list. Temporal grouping — use built-in functions: Good: toStartOfMonth(nli.record_timestamp), toStartOfWeek(...), toStartOfDay(...) Bad: substring(nli.record_timestamp, 1, 7) -- not supported Bad: DATE_TRUNC('month', ...) -- not supported ## String Matching Use CONTAINS for all string matching. Never regex, LOWER(), or UPPER(). Good: WHERE nli.source CONTAINS 'Zendesk' Good: WHERE t.category_enum CONTAINS 'COMPLAINT' For case-insensitive search: (fi.content CONTAINS 'login' OR fi.content CONTAINS 'Login') Exception: when search_graph_values returns an exact canonical taxonomy value, use equality or IN on the taxonomy node name for the pointed query. Good: WHERE l2.name = '<resolved L2 name>' Good: WHERE l3.name IN ['<resolved L3 name A>', '<resolved L3 name B>'] Bad: WHERE l2.name CONTAINS '<resolved L2 name>' -- already discovered exact taxonomy value Bad: WHERE l2.name = '<resolved L2 name>' AND fi.content CONTAINS '<same concept>' ## Entity & Relationship Rules Sentiment vs Category — different entities, different paths: - "negative" / "positive" / "neutral" user intent → Sentiment labels "Negative" / "Positive" / "Neutral" Path: (fi)-[:HAS_SENTIMENT]->(sp:SentimentPrediction) Filter: sp.label CONTAINS 'Negative' CRITICAL: Do NOT query sentiment for Gong/Zoom/Teams/audio sources. These are multi-user conversations — sentiment analysis requires single-user context. If user asks for sentiment on Gong/Zoom: explain the limitation and use category_enum instead. - "complaints" / "praise" / "improvement" / "help" → Category Path: (cft)-[:HAS_THEME]->(t:Theme) Filter: t.category_enum CONTAINS 'COMPLAINT' (or 'IMPROVEMENT', 'PRAISE', 'HELP') Works for ALL sources. Never use sentiment (sp.label) as a substitute for category. Never use t.display_name for category filtering — use t.category_enum. Theme + Subtheme — use separate MATCH clauses from cft, NOT chained paths: Good: MATCH (cft)-[:HAS_THEME]->(t:Theme) MATCH (cft)-[:HAS_SUBTHEME]->(st:Subtheme) Bad: MATCH (t)-[:HAS_SUBTHEME]->(st) -- follows taxonomy hierarchy, not tags Citations: always return nli.record_id (NOT fi.record_id — creates broken links). Inequality: use != (not <>). Compiler rejects <>. Exclusion: use NOT IN ['Gong', 'Zoom'] (not chained != ... AND != ...). Boolean: use OR in a single query, not separate queries. Operator precedence: always parenthesize OR groups when combined with AND. Good: WHERE (a OR b OR c) AND t.type != 'MISC' Bad: WHERE a OR b OR c AND t.type != 'MISC' -- AND only applies to last OR term For taxonomy name filters, prefer exact equality/IN from search_graph_values over CONTAINS. If you must OR taxonomy-name predicates, wrap the entire OR group before adding other filters: Good: WHERE (t.display_name CONTAINS '<term A>' OR t.display_name CONTAINS '<term B>') AND t.type != 'MISC' Bad: WHERE t.display_name CONTAINS '<term A>' OR t.display_name CONTAINS '<term B>' AND t.type != 'MISC' Aliases: use 'total' not 'count' (reserved word). ORDER BY alias name. ## Exploratory Depth (CRITICAL for quality) Two DIFFERENT hierarchies exist — do not confuse them: - L1 → L2 → L3: product taxonomy (feature areas). L3 is a finer category, NOT a complaint. Example: L1="Platform", L2="Integrations", L3="Slack Integration" - Theme → Subtheme: user opinions/complaints about those features. Example: Theme="Slack Sync Delays", Subtheme="Sync fails during peak hours" Feedback is tagged through CustomerFeedbackTags, which links to all five levels: BELONGS_TO_L1, BELONGS_TO_L2, BELONGS_TO_L3, HAS_THEME, and HAS_SUBTHEME. L1/L2/L3 tells you WHICH product feature is discussed. Theme/Subtheme tells you WHAT customers think or experienced about it. When ranking complaints, drivers, requests, praise, or issues, the answer lives at Theme/Subtheme level, not L3. Use L1/L2/L3 as product scope, then join Theme/Subtheme from the same cft for the customer-experience breakdown. Product/feature search with exclusions: - For every included or excluded product/feature concept, first discover its exact taxonomy node and level (L1/L2/L3) using search_graph_values; use search_graph_fields only when field ambiguity remains. - Apply product/feature include and exclude filters through the discovered taxonomy node via BELONGS_TO_L1/L2/L3 and the matching L1/L2/L3 name. - Do not use Theme.display_name or Subtheme.display_name to include/exclude product or feature areas. Use Theme/Subtheme display names only when the user explicitly asks for those opinion themes/subthemes. - Do not use nli.source or content fields unless the user explicitly asks for source/channel/integration filtering. - If no taxonomy node can be mapped, state that taxonomy filtering cannot be applied; do not substitute source, content, or Theme/Subtheme display-name filters. L3 ≠ Theme. L3 tells you WHICH feature. Theme tells you WHAT users say about it. Theme and Subtheme use `display_name` for filters, grouping, returned labels, and presentation. L1/L2/L3 use `name`. Preserve every other object's existing schema-defined field; never switch another object to `display_name` because of this taxonomy rule. "What are complaints?" → needs Theme (t.display_name), NOT L3. "What features exist?" → needs L2/L3. For exploratory "what" questions, you MUST drill to the opinion level: When user asks about complaints, issues, pain points, or frustrated users: - The answer must include t.display_name (specific complaint themes), not just L2/L3 feature names. - If combining with features: return BOTH l2.name AND t.display_name in the same query. - Do NOT run separate queries — one for features, another for themes. Combine them. When user asks about dominant themes or top themes: - First run a Theme-level ranking query only: join (cft)-[:HAS_THEME]->(t:Theme), filter t.type != 'MISC', and aggregate by t.display_name to get exact theme totals. - Do NOT join Subtheme in this first ranking query; subtheme fan-out changes the row grain and can make theme totals look like summed subtheme counts. - Then run a separate subtheme-breakdown query for the returned theme names: use the same filters, add MATCH (cft)-[:HAS_SUBTHEME]->(st:Subtheme), filter st.type != 'MISC', and return t.display_name, st.display_name, and subtheme_total. - Final answers may show a few top subthemes under each theme, but theme totals must come from the first Theme-level query. Never sum subtheme counts to produce theme totals. - Do not use WITH ... ORDER BY ... LIMIT followed by another MATCH in one query. When user asks about issues in a specific area: - Scope to the L2 area, but return t.display_name (theme display names) underneath — not l3.name. When user asks about feature requests or improvements: - Same pattern as complaints but with IMPROVEMENT category filter. Skip extra depth for: counts, sentiment, overview questions. Good: RETURN l2.name AS feature, t.display_name AS complaint, COUNT(DISTINCT nli.record_id) AS total Bad: RETURN l2.name AS feature, COUNT(DISTINCT nli.record_id) AS total (Theme is joined for filtering but t.display_name is missing from RETURN — user only sees feature names) Good first query for top themes: RETURN t.display_name AS theme, COUNT(DISTINCT nli.record_id) AS total Good second query for subthemes: RETURN t.display_name AS theme, st.display_name AS subtheme, COUNT(DISTINCT nli.record_id) AS subtheme_total Bad: RETURN t.display_name AS theme, st.display_name AS subtheme, COUNT(DISTINCT nli.record_id) AS total (Combined Theme+Subtheme rows are subtheme-grain counts, not exact theme totals) Rule: if you join Theme or Subtheme, include their display_name in RETURN unless the query is the first Theme-level ranking query in the top/dominant-theme two-query flow. ## Taxonomy Resolution Prefer taxonomy filters over content keywords when a relevant taxonomy value exists. If the level is unknown, search relevant taxonomy/source types and use the returned object_label/property_path; do not assume the level upfront. Use exact equality/IN on canonical names and filter type != 'MISC' unless miscellaneous is explicitly requested. L1/L2/L3 are product scope. Theme/Subtheme are customer topics/issues. For complaints/topics under a product scope, join Theme/Subtheme from the same cft; complaint requests should filter t.category_enum CONTAINS 'COMPLAINT'. If no relevant taxonomy value exists, use content fallback: summary/content searches use fi.content; exact/raw wording searches use nli.content. Preserve quoted/exact phrases as one string; do not split them into word-level OR terms unless the user asks for a broad keyword search. ## Content Fields fi.content = AI-generated summaries/chapters. Use for summary/content fallback. One NLI → 5+ FeedbackInsights (e.g., Gong call → multiple topical chapters). nli.content = raw verbatim feedback (full transcripts, 50k+ chars). Use for exact/raw wording fallback; do not present returned nli.content as final quote text. ## Quote Workflow This tool is for finding candidate records, not extracting final user quotes. For quote/verbatim/exact-words requests: - Query only candidate metadata needed by find_user_quote: DISTINCT nli.record_id, nli.record_timestamp, nli.source, and fi.content AS summary. - Do not return fi.content or nli.content as a column named quote/verbatim. - Do not render fi.content or nli.content as quoted text in the final answer. - After collecting candidate record_ids, call find_user_quote. If it returns no validated quotes, do not substitute summaries, paraphrases, or raw content. ## Field Types & Operators Null checks: IS NOT NULL / IS NULL (never != null or = null). List fields (ratings, scores, tags) are stored as lists: Good: WHERE nli.appstore_rating CONTAINS '5' Bad: WHERE nli.appstore_rating = 5 -- fails on list fields If search_graph_values is empty or incomplete for a known target value, verify with a field-scoped filter on the user term or a stable substring from it: WHERE nli.{field} CONTAINS '<term>' RETURN DISTINCT nli.{field} LIMIT 20. Use unfiltered RETURN DISTINCT only when the target value is unknown and you need value shape. After mapping to an actual value, pin it with exact = or IN, not broad CONTAINS, unless the field is a list. Source-specific fields (ratings, tags, account names, custom metadata): - These may be omitted from the compact schema. Call search_graph_fields when the exact field is not already available there. - Qualify with the alias matching each field's entity group returned by search_graph_fields: NLI/FeedbackRecord fields use nli.{source}_{field} (e.g. nli.appstore_rating, nli.g2_metadata_starrating); Account/User fields use the joined alias (e.g. acc.industry via PROVIDED_BY_ACCOUNT). - Never guess an exact field name; use the compact schema or search_graph_fields. - For filtering by rating/score: use CONTAINS since values are stored as lists. - Common NLI source fields: nli.appstore_rating, nli.appstore_country_full / nli.appstore_country_full_name, nli.g2_metadata_starrating, nli.gong_account_account_name, nli.gong_account_opportunity_name, nli.gong_metadata_title. ## Supported Syntax - MATCH, OPTIONAL MATCH - WHERE with CONTAINS, =, !=, >, <, >=, <=, IS NULL, IS NOT NULL, IN, NOT IN, AND, OR, NOT - RETURN with aliases, ORDER BY, LIMIT, DISTINCT - COUNT(DISTINCT ...), SUM(), AVG(), MIN(), MAX() - toStartOfMonth/Week/Day(), now(), INTERVAL, toDateTime(), BETWEEN - CASE WHEN ... THEN ... ELSE ... END (prefer WITH for bucketed counts) - HAVING for post-aggregation filtering - COALESCE() for null fallbacks ## OPTIONAL MATCH Use when user asks to INCLUDE items WITHOUT a relationship (e.g., "include uncategorized"): Good: OPTIONAL MATCH (fi)-[:HAS_TAGS]->(cft)-[:BELONGS_TO_L2]->(l2) RETURN COALESCE(l2.name, 'Uncategorized') AS category, COUNT(DISTINCT nli.record_id) AS total Alternative — run TWO separate queries: Query 1: Count by category (normal MATCH with taxonomy) Query 2: Count total items (MATCH without taxonomy join) Then: Uncategorized = total - sum of categorized Both approaches are valid. Use whichever is simpler for the question. FAIL pattern: using only regular MATCH when user explicitly asks for uncategorized items. ## CASE Expressions Use CASE for bucketing/categorizing values. For grouped CASE bucket counts, prefer WITH ... CASE ... END AS bucket, then RETURN bucket, COUNT(DISTINCT ...). Good: WITH nli, CASE WHEN nli.source CONTAINS 'Gong' THEN 'Internal' WHEN nli.source CONTAINS 'Zendesk' THEN 'Support' ELSE 'Other' END AS channel RETURN channel, COUNT(DISTINCT nli.record_id) AS total Good: WITH nli, nps, CASE WHEN nps.value <= 6 THEN 'Detractor' WHEN nps.value <= 8 THEN 'Passive' ELSE 'Promoter' END AS category RETURN category, COUNT(DISTINCT nli.record_id) AS total Alternative: run separate queries per bucket with WHERE filters. Both approaches are valid. ## Not Supported - SKIP / OFFSET → no pagination, use LIMIT only - <> operator → use != instead - Multiple WITH clauses / WITH + COLLECT chains - substring(), toString(), DATE_TRUNC(), LOWER(), UPPER(), regex ## Post-Aggregation Filtering Use HAVING (not WITH...WHERE): Good: RETURN source, COUNT(DISTINCT nli.record_id) AS total HAVING total > 100 Bad: WITH source, COUNT(...) AS total WHERE total > 100 Use HAVING for: "sources with more than 100 items", "themes with at least 50 complaints", "show only high-volume channels", any threshold on aggregated counts. ## Deal/Opportunity Queries - "Deal" = "Opportunity". ALWAYS MATCH (opp:Opportunity), never (d:Deal). - All deal queries START from (opp:Opportunity). Never from DealReport, envelopes, or leaf entities. - Opportunity time-range filters and ordering: use opp.close_date for open, won, lost, and closed deals (NOT nli.record_timestamp). Combine it with the appropriate stage filter. - Expected, scheduled, or planned close windows: use opp.close_date. - "Recently closed" means ALL closed deals (Won + Lost), not just Won. - Stage filtering: opp.stage CONTAINS 'Won' or 'Lost'. For live/open deals, exclude won/lost stages or use the org's active pipeline stage values. Analysis pivot: (opp:Opportunity)-[:IS_ANALYSED_AS]->(dr:DealReport) DealReport branches by prediction type, each via its own envelope: - Win/loss: (dr)-[:HAS_WINLOSS_PREDICTION]->(wle:DealWinLossEnvelope) -[:REFERS_TO_WINLOSS_REASON]->(wlr:WinLossReason) Alt one-hop: (dr)-[:HAS_WIN_LOSS_REASON]->(wlr:WinLossReason). Use wlr.name for reason text. Filter opp.stage CONTAINS 'Won' or 'Lost'. - Competitors: (dr)-[:HAS_COMPETITOR_PREDICTION]->(dce:DealCompetitorEnvelope) -[:REFERS_TO_COMPETITOR]->(c:Competitor) Use c.name for competitor name. - Use Cases: (dr)-[:HAS_USE_CASE_PREDICTION]->(due:DealUseCaseEnvelope) -[:REFERS_TO_USE_CASE]->(uc:UseCase) Use uc.name for use case name. - Features: (dr)-[:HAS_FEATURE_PREDICTION]->(dfe:DealFeatureEnvelope) -[:REFERS_TO_FEATURE]->(df:DealFeature) Filter dfe.mention_type for COMPLAINT/PRAISE. NEVER use L2/L3/Theme/CFT for "features in deals". - Risks: (dr)-[:HAS_RISK_PREDICTION]->(dre:DealRiskEnvelope) -[:REFERS_TO_RISK]->(risk:Risk) Alt semantic hop from the envelope: (dre)-[:HAS_RISK]->(risk:Risk). Use risk.name for risk text. For live-deal risk counts, count DISTINCT opp.record_id. Treat a successful, correctly scoped result from this path as authoritative, including zero rows. Do not add risks inferred from summaries, calls, or feedback. Only if the prediction path is unavailable may you return a separate `LLM-inferred potential risks` section; disclose that fallback and never combine it with structured predictions. For a risk call snippet, first inspect the active schema. The call is linked only when a declared path from the returned DealRiskEnvelope reaches that call's NaturalLanguageInteraction.record_id. Do not guess a relationship from another envelope type. If no such path or row exists, report no linked supporting call moment; do not substitute generic deal feedback, DealReport prose, or quote search. Feedback from deals — two DIFFERENT paths, do not confuse them: - ALL feedback on a deal (any topic): (opp:Opportunity)-[:HAS_FEEDBACK_RECORD]->(nli:NaturalLanguageInteraction) Return fi.content verbatims, not just metadata. Use DISTINCT for samples. - Insight-specific feedback (for an envelope type whose exact active-schema relationship to Evidence has been verified): envelope pivot → leaf filter + (envelope)-[:HAS_EVIDENCE]->(ev:Evidence)-[:SOURCED_FROM]->(nli:NaturalLanguageInteraction) Alt semantic Evidence→NLI hop: (ev)-[:HAS_FEEDBACK]->(nli:NaturalLanguageInteraction). NEVER use Opportunity→HAS_FEEDBACK_RECORD for insight-specific — it returns ALL feedback. Never copy this pattern to DealRiskEnvelope by analogy; follow the risk rule above. ## Out of Scope Decline without executing queries: creating/modifying platform artifacts, exporting files, rendering charts, UI navigation, product how-to questions. Never name platform tools (Quantify, dashboards) unless the user mentioned them first. Never expose internal queries (Cypher, SQL) to the user. Args: cypher_query: Valid Cypher query following the rules above. If the user asked "how many", "all", "total", "count", or "every" — do NOT add a LIMIT clause. description: Plain language description of your information need. Returns: Structured query results with row count, columns, and data. If source-specific metadata is available, also includes relevant_source_fields. Current date for relative date resolution: 2026-09-16 UTC. CITATION URLS: The response carries a ``rows_v2`` field where every COUNT cell is wrapped as ``{"value": <number>, "citation": "<dashboard URL>"}``. When ``rows_v2`` is populated, ``rows`` is empty — use ``rows_v2`` exclusively and wrap the cited number (plus its noun / qualifier when present in prose) inside a markdown link so the URL fires when the reader clicks the claim: - prose: ``[42 complaints](https://dashboard.../citations/...)`` - table cell: ``[42](https://dashboard.../citations/...)`` Don't fabricate URLs; only use ones returned by this tool.

  • Search Graph Values

    Search indexed Context Graph field/value suggestions to find candidate entity, source, taxonomy, and filter values before writing Cypher. Purpose: use this as the broad concept locator to discover where a user phrase appears in the graph and resolve exact values for filters. This tool intentionally returns both field/property matches and value matches because user wording can refer to a taxonomy value, source name, metadata value, product/account/status/segment value, or a field whose values need follow-up. TWO PHASES -- the second one is yours to ask for: - WITHOUT `entity_types`, this searches LITERALLY across every field: metadata, source, and taxonomy. Cheap, broad, and it tells you WHERE a value lives. It cannot match a value worded differently from your term. - WITH `entity_types`, taxonomy values are additionally searched SEMANTICALLY, which is the only way to reach a reordered, inflected, or paraphrased value ("pin calendar sidebar" -> the stored "Want Pinned Sidebar Calendar"). - So after a phase-1 call, CALL AGAIN with `entity_types` whenever the concept is a topic, issue, feature, or theme -- EVEN IF phase 1 already returned values. Phase 1 returning something is not evidence that it returned everything: it finds the literal match and misses the paraphrase of the same idea. - Do NOT escalate when the concept is not a topic and phase 1 already resolved it on a metadata or source field -- a person, account, ticket, email, or source name is not a taxonomy label, and semantic search on those fields does not exist. Take the phase-1 value and move on. - Semantic search exists ONLY for: Theme, Subtheme, L1, L2, L3, Competitor, WinLossReason, UseCase, DealFeature, Risk. Nothing else is semantically searchable, so nothing else is worth escalating. WHICH `entity_types` TO ESCALATE WITH -- a wrong level returns almost nothing: - Name the level the USER meant. "themes" -> ["Theme"]; "subthemes" or specific complaints -> ["Subtheme"]; "product areas"/"categories"/"features" -> the L1/L2/L3 level that matches how this org's taxonomy is organised (check the schema); competitors -> ["Competitor"]. - If the level is genuinely unknown, pass ALL FIVE: entity_types=["Theme", "Subtheme", "L1", "L2", "L3"]. Measured: this finds exactly as much as naming the right level, and returns ~2.8x as many values to read. That is the correct trade when you are unsure -- it costs reading, never recall. - It is ALL FIVE or the level the user named -- there is no middle option. Two or three levels is still guessing, and it carries the same risk as one while saving almost nothing: the value sits at exactly one level, so a partial guess either includes that level or misses it entirely. - NEVER guess. Measured: the right level recalls ~88% of expected values, a wrong level recalls under 1% -- and it does NOT come back empty. Semantic search ranks neighbours with no relevance floor, so the wrong level hands you plausible-looking values that are not what you asked for. A gibberish term still returns 25 of them. Treat a scoped result as an answer only if a value in it is actually the thing you searched for. - Do NOT pick the level from which levels phase 1 happened to match lexically. Measured: that is worse than naming the level AND noisier, because a common word matches several levels while the semantic answer lives in one. - The response names the levels it did NOT search. Read that list before concluding a value does not exist, and before accepting a value you are not sure about. Mandatory workflow: - Before any `run_graph_query` WHERE filter using a user-provided value, alias, source/channel name, taxonomy/product term, segment, status, rating, score, account, competitor, or metadata value, call `search_graph_values` for that concept unless the exact graph value came from tool output earlier in this turn. - For taxonomy/product value searches, use this tool to discover the matching taxonomy field and canonical value. A broad taxonomy search can identify the candidate level; use the returned object_label/property_path for any focused follow-up before writing the final WHERE clause. - For intent phrases such as "why did users leave/churn/cancel/deactivate", search the outcome concept (`queries=["deactivation", "churn", "leave"]`) before querying churn/deactivation reasons. Do not spend the first value search on the customer/org/product name (for example "Nextdoor") unless the user is explicitly asking to filter by that product, source, or account. - LITERAL TERM vs PARAPHRASE decides what this tool's output MEANS. It returns a RANKED, CAPPED candidate list (100 values per field); a field-scoped Cypher scan returns the COMPLETE set. Search first, per the rule above — then read the result according to which case you are in: (a) The user's word appears INSIDE the stored values ("competitor", "calendar", "SSO") AND the answer depends on having every match ("how many themes…", "which subthemes…", "list all…"). Treat what came back as a SAMPLE, not the set, and enumerate before answering: MATCH (st:Subtheme) WHERE st.display_name CONTAINS 'competitor' AND st.type != 'MISC' RETURN DISTINCT st.display_name LIMIT 500 Anchor on the taxonomy node itself, NOT through NaturalLanguageInteraction — an NLI-anchored query reaches only values that already have tagged feedback and silently drops the rest. Theme uses `display_name`, L1/L2/L3 use `name`, and every level needs `type != 'MISC'`. CONTAINS is case-insensitive, so pass the user's term as typed. If the row count EQUALS the LIMIT the set is truncated, not complete — get the real breadth first with `RETURN COUNT(DISTINCT st.display_name)`. This is the one situation where CONTAINS beats the resolved-value filter: you are enumerating the catalog, not filtering records. (b) The user's wording does NOT appear in the stored value — a paraphrase, reordering, or inflection ("pin calendar sidebar" → "Want Pinned Sidebar Calendar"). This tool is the ONLY thing that can resolve these. No CONTAINS predicate will ever match them, so do not fall back to one; if the value came back here, filter on it with exact `=`/`IN`. (c) This tool returned empty or ambiguous — repair with a FIELD-SCOPED fuzzy scan (`WHERE <field> CONTAINS '<term-substring>' RETURN DISTINCT <field> LIMIT 20`), per "WHEN SEARCH RETURNS EMPTY OR INCOMPLETE" below. (d) This tool warned that more values matched than it returned — the (a) recipe is how you get the rest. Never present a capped list as complete. Do NOT open with an UNFILTERED `RETURN DISTINCT <field> LIMIT 20` on a value the user just named. An arbitrary 20-value window is not case (a) — it has no CONTAINS predicate, so it hides the right value below the LIMIT rather than enumerating matches. - If a value search returns empty, do NOT drop the filter, keep the user's spelling, or use the unverified term directly in Cypher. Empty means "not in the indexed suggestion set" (see cap below), not "not in the graph." Follow the 4-step repair in "WHEN SEARCH RETURNS EMPTY OR INCOMPLETE" below. Worked example — Nextdoor user asks about the Qualtrics SMB NPS "spenders" bucket: 1. search_graph_values(queries=["spenders", "spender", "spending"]) → empty 2. Schema / search_graph_fields identifies the likely field: nli.qualtrics_metadata_bucket 3. Field-scoped fuzzy scan with run_graph_query: MATCH (nli:NaturalLanguageInteraction) WHERE nli.qualtrics_metadata_bucket CONTAINS 'spend' RETURN DISTINCT nli.qualtrics_metadata_bucket LIMIT 20 → ['spending', 'active_not_spending', 'not_spending_not_active'] 4. Filter with the verified value using EXACT `=` or `IN` — NOT CONTAINS, because 'spending' is a substring of 'active_not_spending': WHERE nli.qualtrics_metadata_bucket = 'spending' Caveat in the final answer that "spenders" was resolved to the stored value "spending". TAXONOMY VALUE RESOLUTION: Use this tool to locate candidate L1/L2/L3/Theme/Subtheme/Source values and return canonical names for filters. Search all relevant taxonomy nodes when the user's wording does not specify a level. If a result is relevant, use the returned object_label/property_path and matched value for focused follow-up or exact filtering. Once resolved, Cypher should filter taxonomy names with exact equality/IN on the canonical value, not broad CONTAINS predicates. The single exception is case (a) above: when the question is WHICH VALUES EXIST for a literal term, CONTAINS on the taxonomy node is the enumeration — you are listing the catalog, not filtering records against a resolved value. MISC TAXONOMY SEARCH RESULTS: Ignore MISC taxonomy values unless the user explicitly asks for miscellaneous results. Use for: - Vague concepts such as "activation", "paid teams", or "onboarding". - Source names such as "Zendesk", "Gong", or "App Store". - Source/channel aliases such as "support tickets", "call transcripts", "survey", or "app reviews". - Taxonomy/entity labels: Theme/Subtheme display names and L1/L2/L3 names. - Account, product, competitor, work item, status, segment, or metadata values when the user provides the value. Do not use this as the only authority for source/customer metadata field names. For source-specific custom fields such as Zendesk status, App Store rating, Gong title, Salesforce churn reason, plan type, or bucket, use the schema first and `search_graph_fields` when the exact property is unclear. Use explicit type intent when available: - If the user asks for a source/source channel value, pass `entity_types=["NaturalLanguageInteraction"]` on the first call. - If the user explicitly names a Theme, Subtheme, L1, L2, or L3 value, pass that entity type on the first call, for example `entity_types=["Theme"]` or `entity_types=["L1"]`. - If the user gives a product/topic/issue phrase without an explicit level, search the relevant taxonomy levels together to discover where the value is present, then use the returned object_label/property_path for focused verification when needed. - If the user names a source-specific metadata field such as Zendesk status, G2 star rating, plan type, bucket, QID2, or NPS score, resolve the property with the schema/search_graph_fields first, then run field-scoped value verification. Do not rely on a broad value hit from a different field. This is broad discovery over the indexed graph-value suggestion set, not an exhaustive scan of every value in every field. High-cardinality fields may only have a capped sample indexed, so an empty result means "not found in the indexed suggestion set", not "the value does not exist." Without `entity_types`, a search runs a lexical pass over every indexed field plus -- unless the term matched a source name (`NaturalLanguageInteraction.source`) -- a semantic sweep over the taxonomy value fields (`Theme.display_name`, `Subtheme.display_name`, `L1.name`, `L2.name`, `L3.name`) and, for orgs with Sales Insights, `Competitor.name`, `WinLossReason.name`, `UseCase.name`, `DealFeature.name`, `Risk.name`. Sweep values are ranked candidates that may be paraphrases of the term, not literal hits (flagged in `warnings`). Matched values for all other broad fields remain indexed samples; if a needed value is missing, scope the search to the known entity/field or run field-scoped Cypher verification before assuming absence. ONE CONCEPT = ONE CALL with related terms in the queries list: search_graph_values(queries=["ApplePay", "Apple Pay"]) -- spelling variants search_graph_values(queries=["SMB", "smb"]) -- case variants search_graph_values(queries=["paid teams", "paid team"]) -- same value variants DIFFERENT CONCEPTS = SEPARATE PARALLEL CALLS: search_graph_values(queries=["Zendesk", "zendesk"]) -- source search_graph_values(queries=["Gong", "gong"]) -- source search_graph_values(queries=["authentication", "login", "auth"]) -- topic search_graph_values(queries=["spenders", "spender", "spending"]) -- value variants search_graph_values(queries=["SMB", "smb"]) -- segment Never combine unrelated concepts in one call: Bad: search_graph_values(queries=["Zendesk", "NPS", "bucket"]) Bad: search_graph_values(queries=["deactivation reasons churn leaving cancel"]) Related intent synonyms can be one concept when they describe the same outcome: Good: search_graph_values(queries=["deactivation", "churn", "leave"]) Good: search_graph_values(queries=["cancel", "cancellation", "unsubscribe"]) Source resolution: - Always search user-provided source names and source/channel aliases before filtering `nli.source`. - Scope source/source-channel searches to source-bearing feedback entities: `entity_types=["NaturalLanguageInteraction"]`. - Include likely platform variants in the same source concept call, but keep different sources separate. Examples: User: "Show feedback from Zendesk source" Call: search_graph_values( queries=["Zendesk", "zendesk", "ZendeskSupport"], entity_types=["NaturalLanguageInteraction"] ) User: "Compare support ticket channel versus call transcripts" Calls: search_graph_values( queries=["support ticket", "support tickets", "Zendesk", "Intercom", "helpdesk"], entity_types=["NaturalLanguageInteraction"] ) search_graph_values( queries=["call transcript", "call transcripts", "Gong", "call recording"], entity_types=["NaturalLanguageInteraction"] ) Vague metadata values: - When the user gives a vague value such as "paid teams", "free plan", "spenders", "SMB", "enterprise", "active", or "cancelled", call this tool for that value even if the exact field is unknown. - A property-only result can identify a candidate field, but it does not prove the value exists. Follow with schema/field resolution and field-scoped value verification when needed. Field labels versus values: - Use `search_graph_fields` for source/customer metadata field labels such as "bucket", "QID2", "NPS score", "ticket status", or "plan type" when the exact property is unclear. - Still use `search_graph_values` for concrete values in the same user request, such as "Qualtrics", "SMB NPS", "spenders", "Free", or "Canva for Teams". - For mixed source/survey/metadata asks, split the work: search_graph_fields(query="qualtrics bucket") search_graph_fields(query="qualtrics qid2") search_graph_values(queries=["Qualtrics", "qualtrics"], entity_types=["NaturalLanguageInteraction"]) search_graph_values(queries=["SMB NPS", "smb nps", "NPS"]) search_graph_values(queries=["spenders", "spender", "spending"]) - If the value search for a user term is empty, do not drop the filter or accept the user's spelling blindly. Resolve the field, then verify values under that field; for example a user may say "spenders" while the stored value is "spending". If search returns the needed field/value, use that verified value directly in Cypher; do not run fallback Cypher just because the value is important. WHEN SEARCH RETURNS EMPTY OR INCOMPLETE — do NOT conclude the value is absent: 1. Use the in-context schema and/or search_graph_fields to identify the likely field. 2. Run a field-scoped fuzzy match on the user's term (or a stable substring from it) with run_graph_query: MATCH (nli:NaturalLanguageInteraction) WHERE nli.{field} CONTAINS '<term>' [ OR nli.{field} CONTAINS '<Term>' ... ] RETURN DISTINCT nli.{field} LIMIT 20 This scan runs against the graph directly and is not subject to the value-index cap, so it can surface real values missing from indexed samples. Use RETURN DISTINCT nli.{field} LIMIT 20 without WHERE only when the user's term itself is unknown and you need the field's value shape. 3. Map the user's term to the closest actual value from that field-scoped result. 4. Use the verified field/value in Cypher with exact = or IN matching; do not use CONTAINS on the resolved value when sibling values can contain it (e.g. `spending` also appears in `active_not_spending`). Caveat if the value could not be verified. 5. Topic/issue concept where taxonomy search AND the exact-existence check (below) are both empty: the taxonomy may have no label for it, which is not "no feedback about it". Search content directly and say so: MATCH (nli:NaturalLanguageInteraction)-[:SUMMARIZED_BY]->(fi:FeedbackInsight) WHERE fi.content CONTAINS '<term>' OR fi.content CONTAINS '<variant>' Only after taxonomy resolution was attempted -- never as the first step. EXPLICIT TYPE + UNCERTAIN VALUE: If the user names the target type, constrain search to that type and search same-concept variants of the provided value. Do not broaden to other entity types in the same call. Use the named type directly: source/channel -> entity_types=["NaturalLanguageInteraction"] Theme -> entity_types=["Theme"] Subtheme -> entity_types=["Subtheme"] L1/L2/L3 -> entity_types=["L1"] / ["L2"] / ["L3"] source-specific metadata field -> resolve the exact field first, then verify values on that field with field-scoped Cypher. Example: User: search for theme "Terminal activation and app setup issues" Call: search_graph_values( queries=[ "Terminal activation and app setup issues", "terminal activation and app setup issues", "terminal activation", "app setup", "activation", "setup issues" ], entity_types=["Theme"] ) Variant rules: - Keep the exact user phrase first. - Add casing, punctuation, hyphenation, singular/plural, and meaningful substring variants for the same concept. - Do not mix unrelated concepts in one call. - Do not add parent/child taxonomy types when the user specified one type. - If constrained search misses the exact typed value, verify with `run_graph_query` before substituting a nearby result. Result interpretation: - VALUE hit with matched_values: matched_values are candidate exact filter values. - A warning that matched_values "may include semantic (approximate) matches" means they came from the similarity sweep: real stored values, not all of them matches for the user's term. Filter on the ones that fit the wording, not on all of them. Treat other broad matched_values as samples, not a complete field inventory. - Use the returned `object_label`, `property_path`, and `matched_values` when writing Cypher. Do not resolve a value on one field and then silently filter a different field. For example, if "dashboard" resolves to `DealFeature.name`, query the DealFeature path/name rather than guessing `L1.name CONTAINS "Dashboard"`. - If the user explicitly asks through a source-specific field/context (for example Gong account names/titles, G2 star ratings, Zendesk ticket status), a generic metadata value hit such as `uf_brand_size_*`, generic plan fields, taxonomy text, or `fi.content` does not override that field intent. Use search_graph_fields and field-scoped verification on the source-specific field; if no rows match, say no matching source-specific rows were found instead of broadening to a generic value field. - PROPERTY hit without matched_values: this identifies a candidate field/path only. It does not validate that the user value exists under that field. - PROPERTY hit plus matched_values: the field is relevant and the listed values are candidate filter values. - Empty matched_values is not proof of absence unless an exact/exhaustive verification query was run. has_more: - has_more=true means more result entries exist beyond the requested limit. - It does not mean more values exist under the same field. - has_more=false does not prove all values under a returned field were searched. - A warning "<field>: N values shown but more matched" means that field's value list was cut at the per-field cap. Run the Cypher given in the warning (or narrow the term) to see the rest before treating the list as complete. - To see more candidate fields/entities/property paths, call this tool again with a larger limit. - To inspect more values under a resolved field, use `run_graph_query` with `RETURN DISTINCT <field> LIMIT N`. Exact taxonomy verification: If the user provided an exact Theme/Subtheme/L1/L2/L3 phrase and typed search returns empty or only nearby values, do not substitute the nearby value yet. Run an exact existence query on the object-specific label field. Theme and Subtheme use `display_name`; L1, L2, and L3 use `name`. For example: MATCH (t:Theme) WHERE t.display_name = "Terminal activation and app setup issues" RETURN t.record_id, t.display_name LIMIT 5 Use the exact row if it exists. If constrained search and exact verification both fail, the concept most likely has no taxonomy label: fall back to content search (step 5 above) and say so. Ask for clarification only when the term itself is genuinely ambiguous. Args: queries: SHORT terms for ONE concept -- not sentences. One to three words each (spelling, casing, plural, punctuation or substring variants), or a single exact value phrase you already know verbatim. Multi-term calls are sent together so exact values can be ranked against all provided variants. Use a single-element list for the simple case (e.g. ["Zendesk"]). Do NOT invent descriptive sentences as terms. Stored values are labels, not sentences, so a phrase like "interface changes hard to adapt" matches nothing that "interface change" would not, and costs recall by splitting the term across words the label does not contain. The user's phrasing belongs in `criteria`, which is what ranking uses it for: WRONG queries=["password reset email never arrives", "cannot reset my password"] RIGHT queries=["password reset", "reset password"], criteria="password reset emails do not arrive" entity_types: Optional list of entity types to search, such as ["Theme"] or ["L2"]. Use this whenever the user or context implies the value type. For source/source-channel searches, prefer ["NaturalLanguageInteraction"]. limit: Maximum number of result entries per call (default: 10). All terms in `queries` are searched together in one request. criteria: Optional. A SINGLE-TOPIC CRITERION describing what the values should be ABOUT -- the same shape `find_user_quote` takes as `criteria`, e.g. "fees are too high or feel unfair". At most 15 words. This is NOT the user's question, NOT their turn verbatim, and NOT a compressed copy of either. Write it as a condition on the value: "customers want the calendar pinned in the sidebar" "identity, KYC or document verification is failing or blocked" "account is suspended, banned, disabled or cannot be recovered" ONE topic only. Mixing topics, or naming a parent category instead of the concept you are resolving, pulls ranking toward the wrong cluster. Measured, same terms and target each time: criterion on-topic reports 5/8 -> 7/8, account-access 7/8 -> 8/8 parent category ("trust and safety" for verification/KYC terms) verification 14/17 -> 11/17 words lifted from the turn no better than passing nothing DECISIVE for a short or ambiguous term: "calendar issue" alone cannot reach the stored value "Want Pinned Sidebar Calendar" (not found), and lands it at rank 1 with a criterion. OMIT IT when the ask needs BREADTH over many matching values: a criterion concentrates ranking, so on a term with hundreds of valid matches it pushes some past the per-field cap. BAD, and DROPPED with a warning: the full user turn, a system or skill prompt, an uploaded-file notice, a pasted ticket, any URL, anything multi-line -- the generic tokens they carry drag the embedding off the search term. Omit for scheduled/automated callers with no user question. It never changes which fields are searched, only the ranking of candidate values on semantic-search-enabled fields. user_context: Deprecated alias for `criteria`, accepted so existing callers keep working. Prefer `criteria`. If both are sent, `criteria` wins. Returns: Search results with candidate matched entities, property paths, indexed/sample values, matched_values, warnings, and has_more metadata. For entity-filtered taxonomy searches, exact Theme/Subtheme `.display_name` values and L1/L2/L3 `.name` values are preferred filter values for later Cypher. Taxonomy and sales-insight value hits on an unscoped search come from the semantic sweep (see `warnings`); other broad fields are sampled candidates and require scoped follow-up before treating missing values as absent.