The Moz Q&A Forum

    • Forum
    • Questions
    • My Q&A
    • Users
    • Ask the Community

    Welcome to the Q&A Forum

    Browse the forum for helpful insights and fresh discussions about all things SEO.

    1. SEO and Digital Marketing Q&A Forum
    2. Categories
    3. Intermediate & Advanced SEO
    4. Indexing/Sitemap - I must be wrong

    Indexing/Sitemap - I must be wrong

    Intermediate & Advanced SEO
    8 4 199
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as question
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • fretts
      fretts last edited by

      Hi All,

      I would guess that a great number of us new to SEO (or not) share some simple beliefs in relation to Google indexing and Sitemaps, and as such get confused by what Web master tools shows us.

      It would be great if somone with experience/knowledge could clear this up for once and all 🙂

      Common beliefs:

      • Google will crawl your site from the top down, following each link and recursively repeating the process until it bottoms out/becomes cyclic.

      • A Sitemap can be provided that outlines the definitive structure of the site, and is especially useful for links that may not be easily discovered via crawling.

      • In Google’s webmaster tools in the sitemap section the number of pages indexed shows the number of pages in your sitemap that Google considers to be worthwhile indexing.

      • If you place a rel="canonical" tag on every page pointing to the definitive version you will avoid duplicate content and aid Google in its indexing endeavour.

      These preconceptions seem fair, but must be flawed.

      Our site has 1,417 pages as listed in our Sitemap.  Google’s tools tell us there are no issues with this sitemap but a mere 44 are indexed!  We submit 2,716 images (because we create all our own images for products) and a disappointing zero are indexed.

      Under Health->Index status in WM tools, we apparently have 4,169 pages indexed.  I tend to assume these are old pages that now yield a 404 if they are visited.

      It could be that Google’s Indexed quotient of 44 could mean “Pages indexed by virtue of your sitemap, i.e. we didn’t find them by crawling – so thanks for that”, but despite trawling through Google’s help, I don’t really get that feeling.

      This is basic stuff, but I suspect a great number of us struggle to understand the disparity between our expectations and what WM Tools yields, and we go on to either ignore an important problem, or waste time on non-issues.

      Can anyone shine a light on this for once and all?

      If you are interested, our map looks like this :

      http://www.1010direct.com/Sitemap.xml

      Many thanks

      Paul

      1 Reply Last reply Reply Quote 0
      • Martijn_Scheijbeler
        Martijn_Scheijbeler last edited by

        I see your frustration, how long ago did you submit these site maps? Are we talking a couple of weeks or a couple of days/ a day? As I've seen myself, Google is not that fast at calculating the nr of pages indexed (definitely not within GWT). Mostly within a couple of days/ within a week Google largely increased the nr of pages indexed.

        1 Reply Last reply Reply Quote 0
        • SEOAndy
          SEOAndy last edited by

          dealing with your indexing issue first - depending on when you submitted depends how soon those pages may be indexed. I say "may" because a sitemap (yes answering another question) is just an indicator of "i have these pages" it does not mean they will be indexed - indeed unless you've a small website you will never have 100% indexation in my experience.

          Spiders (search robots) index / visit a website / page via another link. They follow links to a page from around the web, or the site itself. The more links from around the web the quicker you will get indexed. (this explains why if you've 10,000 pages you won't ever get a link from other websites to them all and so they won't all get indexed). This means if you've a web page that gets a ton of links it will be indexed sooner than those with just 1 link - assuming all links are equal (which they aren't).

          Spiders are not cyclic in their searching, it's very ad-hoc based on links in your site and other sites linking to you. A spider won't be sent to spider every page on your site - it will do a small amount at a time, this is likely why 44 pages are indexed and not more at this point.

          A sitemap is (as i say) an indicator of pages in your site, the importance of them and when they were updated / created. it's not really a definitive structure - it's more of a reference guide. Think of it as you being the guide on a bus tour of a city, the search engine is your passenger  you are pointing out places of interest and every so often it will see something it wan't to see and get off to look, but it may take many trips to get off at every stop.

          Finally, Canonicals are a great way to clear up duplicate content issues. They aren't 100% successful but they do help - especially if you are using dynamic urls (such as paginating category pages).

          hope that helps

          fretts 1 Reply Last reply Reply Quote 1
          • fretts
            fretts last edited by

            Thanks for the quick responses.

            We had a bit of a URL reshuffle recently to make them a little more informative and to prevent each page URL terminating with "product.aspx".  But that was around a month ago.  Prior to that, we were around 40% indexed for pages (from the sitemap section of WM tools), and always zero for images.

            So given that we clearly have more than 44 pages indexed by Google, what do you think that figure actually means?

            SEOAndy 1 Reply Last reply Reply Quote 0
            • Robert_G
              Robert_G last edited by

              I experienced this issue with sandboxed websites.

              Market your products and in a few months every page should be in Google's index.

              Cheers.

              1 Reply Last reply Reply Quote -2
              • SEOAndy
                SEOAndy @fretts last edited by

                I think that as your sitemap reflect your new urls and this is what the index is based on you are likely to have more indexed from what you say. I  would suggest going to "indexed status" under health of GWT and click total index and ever crawled, this may help clear this up.

                1 Reply Last reply Reply Quote 0
                • fretts
                  fretts @SEOAndy last edited by

                  Thanks Andy,

                  What I dont get, is why Google would index in this way.  I can understand why they would weight the importance of a page based on the number/strength of incoming links but not the decision to index it at all when lead in by a sitemap.

                  I just get a little frustrated when Google offers you seemingly definitive stats only to find they are so vague and mysterious they have little to no value.  We should have 1400+ pages indexed, we clearly have more than 44 indexed ... what on earth does the number 44 relate to?

                  SEOAndy 1 Reply Last reply Reply Quote 0
                  • SEOAndy
                    SEOAndy @fretts last edited by

                    44 relates to the number of pages with the same urls as in your sitemap - it is not everything that is index. Your old site is still indexed and being found, as Google visits those pages and gets redirected to a new page it is likely that number will increase (from 44) and the number of old indexed will decrease.

                    Google doesn't index sites on a one-off go around because then if may take say 4 months to come back and index again and if you've a new important page that gets lots of links and you don't get indexed and ranked for it because you've not been visited you wouldn't be happy. Also if this was done on every site it would take forever and take much more resources than even google has. it is annoying but you've just got to grin and bear it - at least you old site is still ranking and being found.

                    1 Reply Last reply Reply Quote 1
                    • 1 / 1
                    • First post
                      Last post
                    • For a sitemap.html page, does the URL slug have to be /sitemap?
                      ThompsonPaul
                      ThompsonPaul
                      0
                      4
                      205

                    • Client wants to remove mobile URLs from their sitemap to avoid indexing issues. However this will require SEVERAL billing hours. Is having both mobile/desktop URLs in a sitemap really that detrimental to search indexing?
                      RosemaryB
                      RosemaryB
                      0
                      7
                      89

                    • Sitemap Indexation
                      GPainter
                      GPainter
                      0
                      4
                      74

                    • Old/wrong meta-titles in index
                      TimHolmes
                      TimHolmes
                      0
                      3
                      110

                    • Which is better /section/ or section/index.php?
                      TimHolmes
                      TimHolmes
                      0
                      7
                      79

                    • Domaim.com/jobs?location=10 is indexed, so is domain.com/jobs/sheffield
                      rikano
                      rikano
                      0
                      5
                      129

                    • How to remove "/magento/" and "/index.php/" showing in internal links and dup pages in GWT
                      GarGar
                      GarGar
                      0
                      6
                      6.4k

                    • Sitemaps / Google Indexing / Submitted
                      Copstead
                      Copstead
                      0
                      3
                      360

                    Get started with Moz Pro!

                    Unlock the power of advanced SEO tools and data-driven insights.

                    Start my free trial
                    Products
                    • Moz Pro
                    • Moz Local
                    • Moz API
                    • Moz Data
                    • STAT
                    • Product Updates
                    Moz Solutions
                    • SMB Solutions
                    • Agency Solutions
                    • Enterprise Solutions
                    • Digital Marketers
                    Free SEO Tools
                    • Domain Authority Checker
                    • Link Explorer
                    • Keyword Explorer
                    • Competitive Research
                    • Brand Authority Checker
                    • Local Citation Checker
                    • MozBar Extension
                    • MozCast
                    Resources
                    • Blog
                    • SEO Learning Center
                    • Help Hub
                    • Beginner's Guide to SEO
                    • How-to Guides
                    • Moz Academy
                    • API Docs
                    About Moz
                    • About
                    • Team
                    • Careers
                    • Contact
                    Why Moz
                    • Case Studies
                    • Testimonials
                    Get Involved
                    • Become an Affiliate
                    • MozCon
                    • Webinars
                    • Practical Marketer Series
                    • MozPod
                    Connect with us

                    Contact the Help team

                    Join our newsletter
                    Moz logo
                    © 2021 - 2026 SEOMoz, Inc., a Ziff Davis company. All rights reserved. Moz is a registered trademark of SEOMoz, Inc.
                    • Accessibility
                    • Terms of Use
                    • Privacy