A freight tender opens on the first working day of January. The service page that would have won you a place on the bidder list went live in October and has never been fetched by a crawler. Nothing about that page is wrong; it simply does not exist as far as the search engine is concerned.
Discovery is the least discussed part of search work and the easiest to lose money on. It is not glamorous, it produces no chart anybody wants to present, and it fails silently: a page that has never been crawled looks exactly like a page that ranks badly, right down to the zero in the impressions column.
Indexing has a due date you did not set
In most industries a slow crawl is an annoyance. In a port cluster it is a missed window. Contract logistics runs on cycles that begin and end on fixed dates: annual freight tenders, warehousing agreements renewed on an anniversary, framework contracts reviewed once a year, customs representation retendered when a licence changes hands. The buyer's shortlist is compiled in a specific fortnight, and after that fortnight nothing you publish matters until the next cycle.
The same logic applies at the other end of the timescale. A congestion notice, a revised gate schedule, a new hazardous-goods acceptance rule or a temporary depot arrangement has a working life measured in days. If it takes eleven days to be crawled, it was never published in any meaningful sense.
Contract logistics and forwarding
The shortlist is assembled in a known fortnight and closed for another year the moment it is signed.
- Publish and submit a full quarter ahead
- Missing the window costs twelve months, not twelve days
Warehousing and capacity
Enquiries cluster around peak build-up and around the dates when leases and storage agreements come up for renewal.
- Capacity statements need to be current, not merely present
- Last year's figures actively cost you enquiries
Ship repair and marine services
Demand appears with a vessel and disappears with it, so a page that is invisible this fortnight has simply missed the call.
- Availability pages age within days
- Speed of discovery outweighs polish
Chemical and bulk handling
Acceptance rules, classifications and licence conditions change on dates set by somebody else entirely.
- The superseded version must stop ranking, quickly
- Buyers search the rule, not your company name
A logistics site generates more addresses than anybody intended
Sites in this sector are rarely small, and they are almost never as small as their owners believe. The reason is combinatorial. A forwarder with twelve services, forty country pairs and nine cargo categories has not built a hundred pages; the templating system has quietly produced several thousand addresses, most of which nobody has ever read.
Add a filterable depot or equipment list, a document library, a sortable tariff table and two languages, and the crawler is handed a space it cannot finish exploring. It does not announce that it has given up. It simply stops somewhere, and whatever sits past that point stays unfetched.
Filters that became addresses
Every combination of cargo type, service and region produces a distinct URL, and the crawler treats each as a new document worth checking.
- Sort orders and page numbers multiply the same list
- Most combinations return the same handful of results
Parameters bolted on by other teams
Campaign tags, session identifiers and referral codes appended by sales, events or a partner portal, each generating a fresh address for identical content.
- Often introduced without the web team's knowledge
- Frequently linked from external sites, so they get discovered
Two languages, unevenly maintained
A Dutch page and its English counterpart, plus the versions where only one of the two was ever updated after a service change.
- Orphaned translations linger long after the original moved
- Language paths double the crawl surface at a stroke
The document library
Tariff sheets, general conditions, certificates and capacity statements, often published as files with no page linking to them.
- Superseded versions are rarely removed
- Buyers search for these directly and find the wrong edition
None of this is negligence. It is what happens when a site accumulates over fifteen years of service changes, acquisitions and rebuilds, with a document library bolted on somewhere around year six.
Crawl budget is finite and mostly spent badly
A search engine allocates each host a rough amount of fetching per day, based on how quickly your server responds and how much it thinks your content is worth revisiting. You cannot buy more of it. You can only stop wasting the amount you already have.
Waste is the normal condition. The crawler works through variant after variant of the same filtered list while the four service pages you actually want in front of a tender committee wait behind them. Nobody notices, because there is no alert for it.
| What consumes the budget | Typical share on a logistics site | What it should be doing instead |
|---|---|---|
| Filtered and sorted list variants | Large, and growing with every new filter | Fetching the service pages behind them |
| Parameterised duplicates | Whatever sales and events have added | Nothing at all; these need never be fetched |
| Superseded documents | Small in count, heavy in bytes | Fetching the current edition instead |
| Redirect chains from old structures | Grows with every site rebuild | One hop to the live address |
| Slow-responding dynamic pages | Disproportionate, because time is the real currency | Being cached or made static |
| Genuinely new commercial pages | Often a rounding error | Everything |
The fix is subtraction before addition. Removing a few thousand pointless addresses from a crawler's path does more for the pages that matter than any amount of submission, because it changes what the remaining budget is spent on rather than asking for more.
The sitemap is a statement about what you stand behind
A sitemap is often treated as a technical formality generated by a plugin and never looked at again. It is better understood as a declaration: these are the addresses I consider current, canonical and worth your time. When it contains retired lane pages, both language variants of a document that only exists in one, and every filter combination the templating engine can produce, it is not a declaration of anything.
Within the indexing tools in the Semalt panel, sitemaps can be submitted as an uploaded file or as a URL, and nested index files are parsed recursively down to three levels, which matters for the large multi-brand structures common in this sector. A single job handles up to 1,000 sitemaps, two jobs may run at once, and up to twenty can sit queued behind them.
The URL tracker and the sitemap queue
For a site whose commercial pages are outnumbered several times over by machine-generated addresses.
- A daily allowance, not a queue you can flood. One thousand URLs a day per account, with bulk submission of up to ten thousand in a batch, which forces you to decide what genuinely deserves the attention.
- Recursive sitemap parsing. Index files pointing at index files are followed three levels deep, so a group that publishes one sitemap per brand does not have to flatten anything by hand.
- Submission through IndexNow. The API notifies participating crawlers, GoogleBot and BingBot among them, that a given address has changed and is worth refetching.
- A log entry per address. Bot visits with timestamps, the status returned, and the error detail when one comes back, alongside live counters for submitted, discovered and failed URLs.
The daily ceiling is worth taking seriously rather than resenting. A thousand addresses a day is generous for a company with sixty pages that matter and restrictive for one with forty thousand machine-generated ones, and which of those you are is a decision about your site rather than a fact about the tool.
Submission is a request, and nothing more
This is the point at which indexing tools are routinely misrepresented, so it is worth stating without any softening.
That distinction changes how you read the live submission counters. A submitted count is a measure of your own activity. A discovered count is a measure of the crawler's response. A failed count is the only one that reliably tells you something is broken on your side, and it is the one worth reading first every morning.
- Submitted but never visited. The request was accepted and the crawler has not come. Usually a budget problem rather than a page problem, which means the answer lies in what else is competing for the same fetches.
- Visited, then nothing. The bot arrived and the page still does not appear. The crawler judged the content not worth keeping — thin, duplicated, or indistinguishable from four other pages you also publish.
- Visited with an error returned. The most actionable case of the three. A timeout, a server error, an unexpected redirect: a defect with an address attached, fixable this afternoon.
- Indexed once, quietly dropped later. Common on pages that never earn a visit. Recrawling costs the engine something, and it eventually stops paying for a document nobody reads.
Files, editions and addresses that must not move
This sector publishes more downloadable material than most: tariff sheets, standard trading conditions, insurance certificates, hazardous-goods acceptance lists, capacity statements. Buyers search for these directly, often by document name rather than by company, and a broker chasing an acceptance rule at half past six will land on a file rather than a page.
Two failures recur. The first is the superseded edition that stays online and outranks its replacement, usually because it has been linked to for years and its successor was published at a new address with nothing pointing at it. The second is the file that no page links to at all, which will be discovered eventually or never, depending on whether it appears in the sitemap.
There is a procurement dimension here too. When a coordinator forwards your page into a sourcing conversation, that link may be opened weeks later by somebody else, and it will be opened again during the next review cycle. An address that has changed twice in the interval turns into a dead end at the precise moment a stranger is forming an opinion of your company.
| Asset | Belongs in the sitemap | Priority for submission | Common failure |
|---|---|---|---|
| Service and lane pages | Yes, always | Highest, especially before a tender cycle | Buried under filter variants |
| Current tariffs and conditions | Yes | High, and again at every revision | Old edition still indexed and ranking |
| Certificates and licences | Yes | Medium, but check the expiry date | No page links to the file |
| Filtered list variants | No | None | Included by default by the generator |
| Time-limited notices | Yes, while valid | Immediate, then retire deliberately | Left online long after they mislead |
A discovery routine that survives a busy quarter
None of this works as a project. It works as a short recurring habit, ideally owned by whoever also owns the tender calendar, because the two are the same schedule seen from different sides.
Twenty minutes a week, plus one hour a quarter
Enough to keep a mid-sized forwarding or warehousing site discoverable without a dedicated technical role.
- Read the failed counter first. Errors with an address attached are the only part of the report that is unambiguously yours to fix, and they take minutes rather than strategy.
- Submit what changed, not what exists. Revised pages, new documents, retired notices. Resubmitting an unchanged catalogue every week teaches the crawler nothing and consumes the allowance.
- Work backwards from the tender calendar. Anything that must be visible in January is published and submitted in October, so there is time for the fetch, the judgement and the ranking to happen in sequence.
- Prune once a quarter. Retire superseded documents, collapse redirect chains to one hop, and remove filter variants from the manifest. Subtraction is the highest-yield hour in the whole routine.
Where indexing behaviour needs interpreting rather than merely recording, the project assistant in the panel is bound to the actual data: a router decides per question which blocks to load, between none and three of them, replies stream token by token, and up to twenty messages of context are kept. It also accepts pasted URL lists in bulk, which is a faster way to queue a retired-document sweep than clicking through a table. The assistant and its project feed also carry campaign news, new placements and to-dos in one chronological stream.
The two campaign tiers sit alongside all of this and do not replace it: at 149 USD a month per domain the automated tier handles keyword discovery and link building, while the 500 USD tier adds manual keyword selection, placement against a domain-rating target and human review before on-site edits go live. Neither buys you a crawl budget. If your commercial pages are competing with forty thousand filter variants for the crawler's attention, that has to be dealt with first, whatever else you are paying for.
Questions this raises
We submitted a page three weeks ago and it is still not showing. What now?
Check the per-URL log before changing anything. If no bot visit is recorded, the problem is discovery and competition for crawl budget. If a visit is recorded with an error, fix the error. If a visit is recorded and came back clean, the page was fetched and judged not worth keeping, which is a content problem and not an indexing one.
Is a thousand URLs a day enough for a site with tens of thousands of pages?
It is enough for the pages that earn revenue, which on most logistics sites is a small fraction of the total. If the allowance feels tight, that is usually a signal about how many addresses your templating produces rather than about the limit. Bulk batches of ten thousand exist for genuine one-off migrations.
Do we need separate sitemaps for the Dutch and English sections?
Separate files are easier to reason about and easier to audit, and nested index files are parsed three levels deep, so splitting costs nothing structurally. The real benefit is that a gap becomes visible: if one language file lists two hundred addresses and the other lists ninety, you have found something worth investigating.
Our old tariff PDF outranks the current one. How do we fix that?
Redirect the old address to the new file rather than deleting it, so the years of links pointing at it carry over. Then submit the new address and let the log confirm the fetch. Deleting outright loses the accumulated signals and leaves a dead link in every buyer's archive.
Does IndexNow work for search engines that do not participate?
No. It notifies participating crawlers, GoogleBot and BingBot among them, and has no effect on anything outside that set. It is a faster notification channel, not a universal one, and it makes no promise about what happens after the notification is received.
Where to start this week
Begin with a count. Take the number of addresses your site can generate, then the number your commercial team would recognise as a real page. The ratio between those two figures is the whole problem in one line, and on a mature logistics site it is routinely worse than fifty to one.
Then take the tender and renewal dates for your five largest customer categories and mark the publication deadline three months ahead of each. That calendar, not a reporting cycle, is what discovery work is actually scheduled against. How the resulting figures are read once pages are live is covered elsewhere in the English article index, and the structural work belongs with our services.
After that it becomes routine: subtract, declare, submit, read the log. If you would rather watch it happen on your own site than take it on description, open the panel and put your sitemap through the queue, then compare the discovered count against what you believed you had published. The gap between those two numbers is usually the most useful thing you will learn all quarter.
Ready to Improve Your SEO?
Get a free SEO audit and find out how we can grow your Rotterdam business.
Request a Free SEO Audit