On the surface, an e-commerce site search can seem to work perfectly: the search bar is visible, typing an exact product name brings up the right item, and nothing looks obviously broken. But that’s the easiest test there is. It doesn’t tell you whether the site can actually understand its customers.
Shoppers rarely know the exact catalogue name, SKU code, or taxonomy your business uses internally. More often, they search by feature, size, material, or need: “reading lamp with warm light”, “round dining table for six”, “black outdoor wall light”. If your search engine only matches literal keywords, products that are in stock – and potentially a perfect fit – can stay invisible.
You don’t need a complex technical project to find out where you stand – just pen and paper, or a simple spreadsheet, to log what you find. Run your site search through twenty realistic queries and watch not only whether it returns results, but whether it leads shoppers to the right ones quickly. It’s a simple, repeatable check – and it reveals far more than you’d expect.
At a glance
- The “working” search illusion: If your testing only covers exact product names and SKUs, your site search may seem efficient while continuously hiding products from shoppers searching by need, attribute, or natural language.
- The cost of failed queries: Irrelevant results and false “no results found” pages disrupt high-intent customer journeys, leaving shoppers with the impression that what they’re looking for simply doesn’t exist.
- The 20-query test solution: A structured set of realistic searches – evaluated for relevance, clarity, and filtering options – identifies priority issues and establishes a measurable baseline for improvement.
Why site search deserves its own test
Anyone using the search bar is declaring an intention. They aren’t just browsing the site – they’re trying to find something specific. That’s why an imprecise search can do more damage than imperfect navigation. The customer makes a request, the site appears to respond, yet shows irrelevant products or an empty results page.
A study commissioned by Google Cloud and carried out by The Harris Poll, covering almost 13,500 adults across 14 countries, shows just how commercial that intent really is. In the US sample, 69% use site search to find products – and when that search succeeds, 92% go on to buy the item they were looking for, while 78% add at least one more product to their basket.
Baymard’s Search UX benchmark shows that many e-commerce sites still struggle with everyday queries – searches by product type, feature, or even exact code. In lighting and furniture, this is made harder by the sheer number of finishes, sizes, variants, and technical specs involved. A single word might refer to a product family, a function, a shape, or a feature.
So the test shouldn’t only ask “does the search engine find something?” – it should ask “would the customer immediately know they’ve landed in the right place?”
The 20-query test for a lighting or furniture e-commerce site
The test consists of 20 distinct queries, divided into five categories of four searches each. To keep this readable, the examples are grouped into a few points – but each query listed should be run and evaluated separately. Adapt the brands, materials, dimensions, and applications to whichever matter most for your catalogue.
1–4. Name, code and identifiers
This pair checks both basic indexing and tolerance for formatting variations: run two separate searches, first the full name of a product, then a recognisable fragment of that same name. Search for the exact SKU, then repeat the test writing the same code with different spacing, hyphens, or capitalisation.
5–8. Product type and synonyms
Next, check whether your search engine connects customer language to catalogue language: run two separate searches, a generic product type such as “pendant lights”, and a common synonym such as “chandelier”. Then search using the singular instead of the plural, and try a commercial term different from your category name.
9–12. Attributes, variants and combinations
To see how the search engine combines products with characteristics, try three separate queries: product plus colour (“black wall light”), product plus material (“oak table”), and product plus dimension (“60cm diameter pendant”). Finish with a fourth query that combines multiple attributes: “red dimmable 2700K lamp”.
13–16. Typos, abbreviations and units of measurement
This group tests how well your search tolerates different ways of writing the same information: try two separate queries, one word containing a typo, and a name with no accent or a simplified spelling. Then add two more – a measurement in a format different from your catalogue’s (choosing between “60cm”, “60 cm”, or “600 mm”), and a technical abbreviation such as “CRI”, “IP”, or “K”.
17–20. Need, room and context of use
Finally, test whether the engine understands intent beyond the product name: try two separate queries, “lamp for reading on the sofa” and “light for a dining table”. Finish the test with “a chair suitable for a restaurant” and “lighting for a small bathroom”, evaluating each separately.
How to score your results
For each query, assign a score from zero to two across three areas: relevance of the top results, clarity of the results page, and ease of refining the choice.
Use this scale:
- 0 – No results, or results that are clearly wrong.
- 1 – Some useful products, but mixed in with irrelevant suggestions or hard to filter.
- 2 – Relevant results, explained by visible attributes and easy to narrow down.
The maximum score is 120: 20 queries multiplied by three evaluation areas, multiplied by a maximum of two points each. The absolute number matters less than the detail: note which types of query fail, and which attributes come up again and again in the problems. A search that finds SKUs perfectly but doesn’t understand materials, applications, or dimensions works for people who already know your catalogue – not for those who still need to choose.
What might be behind a failed search
The engine isn’t always solely to blame. A query can fail because an attribute is only mentioned in the description, because equivalent values aren’t normalised, because variants aren’t linked correctly, or because the product is missing information about its context of use.
So the test also becomes a check on your information quality. If “reading lamp” returns nothing, you need to find out whether the search isn’t recognising the intent, or whether no product is described as suitable for reading. The fix might need configuration, data enrichment, or both.
From test to improvement plan
Rank the problems by impact. Start with frequent queries that return zero results, then searches that show the wrong products, and finally difficulties with refinement. Keep the twenty queries as a control set: after every change, run the same test again.
To explore this phenomenon and its commercial consequences further, you can read our article on search abandonment. The next step is connecting your test results to real data: your most-used queries, exit rate after search, zero-result searches, and conversions generated by users who search.
Frequently asked questions about e-commerce site search
Is this test still useful if our site search almost always returns results?
Yes, because a full results page doesn’t mean the answer is correct. Customers might see dozens of results, but if the most relevant product isn’t near the top, if the wrong variant shows up, or if there aren’t enough filters to narrow the choice, the experience still breaks down. So the test measures the quality of the response, not just whether results appear at all: relevance, clarity, and how easily someone reaches the right configuration.
How do I select the 20 test queries if I don’t have search analytics yet?
Start with questions gathered from your sales team, customer care, and your dealer network, along with terms found in customer emails and the gaps between your catalogue language and the language used in the market. Include strategic products, attributes that determine compatibility, plausible typos, and at least four queries phrased around a room or a need. Once real data becomes available, gradually replace your assumptions with your most frequent queries.
If the test uncovers issues, should we replace our search engine or fix our catalogue?
You can’t decide that from the number of errors alone. You need to check whether the necessary information exists, is normalised, and reaches the search index. If the data is correct but the engine can’t combine attributes, handle synonyms, or understand context, the fix is technological. If attributes and variant relationships are missing, changing engines will just move the problem, not solve it. In practice, the improvement plan usually touches both areas.
Want to know what your site search is hiding?
At MON-KEY, we analyse queries, results, filters, and catalogue structure to find exactly where the path to the product breaks down. Contact us for an evaluation of your store’s internal search.