Transparency
How we test
How a tool ends up on our list, which five criteria we score it against, how the overall rating is calculated – and where our tests hit their limits. No marketing incantations, just numbers you can recalculate yourself.
How a tool ends up on our list
We don't test everything that exists – we test what is genuinely relevant for e-commerce founders in German-speaking Europe. A tool makes the list if at least one of these applies:
- It solves a problem readers keep telling us about – or one we run into in our own business.
- It is actually usable in the DACH region: German interface, EU hosting, GDPR-ready contracts or German-speaking support.
- It shows up in our audience – in forums, communities, search queries – and there is no useful independent write-up yet.
- It is affordable for beginners or offers a meaningful free tier.
We do not charge for a tool to be tested and we do not sell review slots. Vendor requests are treated like any other suggestion: they can get a tool onto our list – they do not influence what ends up in the article.
The scoring framework: five criteria
Every rated tool is scored across five dimensions, each from 0 to 5 points. We adapt the exact labels to the tool category – a search product gets „Search quality & AI features“, an accounting tool gets „Feature set & DATEV“. Behind the labels, the same five questions always apply:
| Dimension | The question behind it |
|---|---|
| Feature set | Can the tool do what you're buying it for – and what is missing compared to its direct alternatives? |
| Usability & setup | How long until the first meaningful result? Can someone without a technical background handle it? |
| Value for money | What does the realistic entry plan cost – and what happens to that price as the shop grows? |
| DACH & GDPR fit | German interface, EU hosting, DPA, German invoicing and tax logic, German-speaking support? |
| Support & track record | How responsive is support, and how solid is the evidence base from independent sources? |
How the overall rating is calculated
The overall rating is the unweighted average of the five individual scores, rounded to one decimal place. No hidden point system, no weighting we don't show you. You can recalculate every rating we publish – here using our Doofinder review as an example:
Setup & usability 4.8
Platform coverage (DACH) 4.7
Value for money 4.0
Support & reliability 3.6
─────────────────────────────
(4.7 + 4.8 + 4.7 + 4.0 + 3.6) ÷ 5 = 4.36 → overall 4.4
Next to every overall rating we add one sentence explaining where the points were lost. If you're short on time, that is the single most useful sentence in the whole review – it tells you what will annoy you about the tool.
What our ratings don't tell you
Our overall ratings sit between 3.9 and 4.7. That looks suspiciously generous, and it's a fair objection – so here's the explanation: it's a consequence of selection. We write about tools we consider broadly usable. A tool that fails our test doesn't get a bad review; it gets no review at all.
Which means the real differentiation is not in the overall rating but in the five individual scores. Those range from 2.8 to 4.9. A tool rated 4.4 overall can score 3.2 on pricing transparency – and that is exactly the information you came for.
How to read our reviews: look at the lowest individual score first, then at the sentence explaining the deduction. The overall rating is orientation, not a verdict.
Where our data comes from
Honestly: we can't run every tool in live production for months. So we're explicit about what each rating is based on – it's always a mix of these sources:
- Our own test. Create an account, set it up, run the core use case on a real case. This is the basis for the usability, setup and feature-set scores.
- Vendor information. Pricing pages, documentation, feature lists, support hours. We take these as vendor claims – and label them as such.
- Independent review platforms. Trustpilot, G2, Capterra, the app stores of the shop systems, plus reports from forums and Reddit. These mainly feed „Support & track record“, because a single test says little about that.
Where our tests hit their limits
Three constraints we won't argue away, so you can judge the result properly:
- We typically test for days to weeks, not years. Long-term issues – creeping price increases, declining support – often only reach us through third-party sources.
- We test from the perspective of a small to mid-sized shop. Whether a tool still holds up at six-figure order volumes is usually beyond what we can verify.
- Software changes. Every review carries a last-updated date; whatever happened after that only appears with the next update.
Independence & affiliate links
We partly finance this site through affiliate links. If you buy through one of them, we earn a commission – at no extra cost to you. To keep that verifiable, these rules apply:
- Affiliate links are marked with an asterisk (*) and a note in the article.
- Commission rates are not a criterion in the scoring framework – they simply don't appear in it.
- We also recommend tools we earn nothing from when they are the better choice – and we write down weaknesses even when there's a partner programme behind the tool.
- No vendor gets to approve an article before publication.
We're not claiming this makes us automatically neutral – a business model always exerts pressure. But you can check whether we stick to it: the low individual scores in our reviews land almost exclusively on tools with partner programmes.
Found a mistake?
If a price is out of date, a feature has disappeared, or your day-to-day experience with a tool differs from ours: tell us. We correct reviews and reset the last-updated date – that happens regularly, and we much prefer it to an article quietly going stale.
Who does the testing
This page was last reviewed in August 2026. We update it whenever our process changes.