A URL can be live, load perfectly in your browser and still be absent from Google. When that happens, I would not start by submitting it to Search Console over and over. I would first work out where the path has failed.
Google has to discover the URL, crawl it, and decide whether it belongs in its index. A sitemap helps with discovery. Accessible HTML helps with crawling. Neither one forces a page into search results.
This portfolio has a foundation intended to support that path. Astro builds HTML for every route, the sitemap reads posts and projects from the collections, and the SEO component creates canonicals and language alternates. That is a sound production setup, not an indexing guarantee. If a new URL went missing, this is the order I would follow.
I start with the exact URL, not a broad search
The site:example.com operator can be a useful hint, but it is not a diagnostic report. It may show stale results, miss a URL Google knows about, or return a different variant from the one I am investigating.
I would open Search Console’s URL Inspection tool and paste the exact canonical address, including the trailing slash if that is the published version. I would look for three specific answers:
- Does Google know this URL?
- Could it fetch it without a problem?
- What reason does it give for not indexing it, or which canonical did it choose?
I would not treat that as a standalone verdict. Inspection reports what Google currently has for that URL and gives me a much stronger starting point than changing tags at random. After a fix, it also lets me test the deployed page rather than the HTML open on my machine.
I confirm which version should exist
Before looking at Google, I would make sure I am not chasing the wrong URL. It is easy to do after redirects, query parameters, a migration, or when a site has language versions.
This site uses trailing slashes on public URLs. /blog/my-post/ is the version that should appear in the sitemap, internal links and canonical tag. If a route redirects, I would inspect the final destination. A new page that I expect to rank should return a 200 there and contain the real content.
I would also check the HTML that reached production. A post can exist as MDX in a repository and still be missing from a build because of a collection error, a malformed date, or a condition that leaves it out of a listing. On a client-rendered site, I would inspect the initial response as well as what appears after JavaScript runs in my browser.
I check what Google can actually read
The next question is less glamorous and often solves the real problem. What does a crawler receive when it requests that URL?
I would expect a 200 for the page I want indexed. A 404, 410, 5xx, or a chain of redirects makes this a different diagnosis. I would also look for a response that resembles an empty page, an error page returning 200, or a URL that is almost identical to another one. Those cases can look like a soft 404 or duplicate content to Google.
Then I would look in two places for an instruction not to index:
- A
<meta name="robots" content="noindex">tag in the HTML. - An
X-Robots-Tag: noindexHTTP header.
Searching the source code for noindex is not enough. It needs checking in the deployed response because middleware, hosting configuration, or a preview rule can add headers too. Google’s robots meta and X-Robots-Tag documentation highlights another common trap. Google needs to crawl a page to read the directive. Blocking a URL in robots.txt is not a reliable way to remove an already indexed page and may stop Google from seeing a new noindex.
A canonical is not an order Google has to follow
A canonical tells Google which URL I consider primary when several versions are similar. Adding a rel="canonical" tag does not make a page unique.
I would check that the absolute canonical points back to the page itself when the content is genuinely distinct. If a project page accidentally points to the home page, or every paginated page shares one canonical, the site is sending conflicting signals. Search Console can also show a Google-selected canonical that differs from the declared one.
On a bilingual site, I would add one more check. hreflang connects Spanish and English equivalents, but it does not replace a canonical. Every indexable version needs its own canonical and the alternates need to form a consistent group. In my Astro hreflang implementation, I show how I keep that link between routes and sitemap entries.
Google’s guidance on consolidating duplicate URLs is useful here because it frames a canonical as a signal, not an absolute command. If Google picks another URL, I would look for real duplicates, inconsistent links, or redirects that explain the choice before trying to force it.
Sitemaps and internal links do different jobs
This project’s sitemap is generated from static pages, posts and projects. Every new entry is included with its language versions. That gives Google a way to discover URLs which may not have many links yet.
I would not use it as an approval list, though. A sitemap says, “this URL exists and matters to me.” It does not explain why a reader should reach it or why Google should keep it indexed.
I would therefore check both:
- The final URL appears in the published sitemap and is not a redirect or an old variant.
- It receives internal links from pages covering a related subject.
This post, for example, is linked from the portfolio’s technical SEO article and a specific hreflang explanation. Those links are not a trick to push a URL. They provide context and create a real navigation path. When a page only exists in the sitemap and no one can reach it by moving through the site, I would first ask whether its place in the architecture makes sense.
The page may still lack a reason to stay
This is where many technical reviews stop too early. A page can return 200, have no noindex, appear in the sitemap, and still offer too little distinct value.
I would not solve that by padding it with definitions found in a hundred other articles. I would check intent instead. Does the URL answer a specific question? Does it add an example, decision, data, or explanation that is not repeated in another page on the same site? Do its title and content make the same promise?
My blog starts from decisions that already exist in the project. This post can discuss static HTML, canonicals, the sitemap and language versions because those pieces are implemented here. That does not make Google obliged to index it. It does keep me from publishing a generic guide that could belong to any site and prove nothing.
If inspection reports a duplicate page or Google has crawled it without indexing it, I would compare that URL with its internal competitors before changing the sitemap. The right fix can be merging two posts, redirecting an old version, or rewriting the page around a more precise need.
Request indexing at the end
Once I have fixed the cause, I would use URL Inspection to request a crawl for a priority page. I would do it once and record what changed. Repeating the same request for several days does not speed up the process.
Google says crawling may take days or weeks, and a request does not guarantee an immediate appearance or even inclusion in results. Its systems prioritise useful, high-quality content. The official guide on asking Google to recrawl a URL also distinguishes requests for a few URLs from submitting a sitemap for many.
That distinction changes the order of work. When I have published one important post, I inspect the URL, fix what I find, then request a crawl. After a migration or a large release, I check the sitemap and the architecture before going URL by URL.
What I would call a finished review
I would not close the case when I click “Request indexing”. I would record the canonical URL, HTTP status, robots directives, Google-selected canonical, sitemap presence and the internal links that lead to it. Then I would give Search Console enough time to show whether the state changes and which queries it starts to receive.
That record avoids two common mistakes. The first is crediting the last tag we changed when the actual improvement came from a redirect or an internal link. The second is turning every slow-to-index URL into a technical emergency when it may not have a clear job yet.
The useful part of technical SEO is not collecting checks. It is ruling out causes in an order that lets you make a meaningful change. For the foundation behind those checks, see how I approached technical SEO for this Astro portfolio.