I routinely do audits on sites with tens of thousands or hundreds of thousands of pages. We've found good results simply by appending the page info to page Title & Description so:
Designer Handbags | Unique Pocketbooks | Page 3
Welcome to the Q&A Forum
Browse the forum for helpful insights and fresh discussions about all things SEO.
I routinely do audits on sites with tens of thousands or hundreds of thousands of pages. We've found good results simply by appending the page info to page Title & Description so:
Designer Handbags | Unique Pocketbooks | Page 3
It's highly doubtful that just six links trigger the slap. In the situation with my client, it was thousands. So it'll be interesting to learn if this one bounces back or not, or if something else is the cause.
As for why Google would slap for over-saturation of a particular phrase, I believe it was at SMX Advanced last year when I recall Matt Cutts stating that Google was working on surgical implementation rather than broad entire-site impact.
If it does NOT bounce back, check to see if you've got too many inbound links targeting the phrase in question. I've seen a site get surgically whacked this way. Fortunately the site owner controlled many of the source sites and as soon as they changed those anchors, they bounced back.
It's really a process of experimenting over time to find out the method that results in the most URLs indexed that in turn brings the most relevant traffic. Personally I wouldn't have one for each category, yet without tests there's no conclusive reasoning either way.
have you viewed the source of those pages? Looked for rogue links? Or URLs embedded in scripts you thought were hidden from search bots? There are all sorts of reasons for the reading that becomes clear when examining the source view (seen as Googlebot sees it).
Jason,
One option would be to split up the home page. Devote the above-the-fold portion to your primary purpose, then below that, have article snippets. Alternately, split it down the middle - left/right, or some other combination of a grid layout. The other advantage of snippets on the home page is more people (not all but some) would find that additional content.
By putting it below the fold you are less likely to detract from your call to action if that's your home page's purpose, while still ensuring fresh content, and those snippet links count for a lot since Google puts more emphasis on links coming from the home page content area.
Canonical is most effective in reducing the number of pages indexed. Doesn't address any issues related to unique content to the degree you need to add a lot more unique content to each page you do want indexed.
Right that's what I'm concerned about. Needing to be extremely careful in how the work is done. I pity the person having to be responsible for that if they're not aware of these issues. That's all I was trying to get across.
Beth,
Unfortunately when you've got a lot of products where the only difference between SKUs is color and item number, it means you need to artificially implement uniqueness. This requires becoming creative with product names and descriptions. Since I see you've already got product names covered fairly well, let's talk about descriptions.
If you go to a sample page, such as for the Charlotte Cut Beads, look at the page both from a "length of description" aspect as well as an "overall content" aspect.
That very short description is quite likely to be seen as at least partly, if not mostly duplicate from other descriptions on the site. And at the same time, the total weight of unique content compared the the entirety of the page (with the header, side navigation and footer being replicated across many pages on the site) and you have a bigger problem in hoping to get this designated as a unique page.
Even though it may seem like a burden to have to produce paragraphs of content for a single product page, without that extended unique description for every product, you'll always have problems.
You may be better off changing how your site works than trying to struggle with the current issues.
For example, when I go to the Charlotte Cuts category page, maybe you should have this be the single page people go to so they can choose which products to add to their cart - eliminating the final individual product "details" page (since there really are no "details" on those.
Then, I'd recommend adding a couple paragraphs of well written and optimized content for that category page.
Agreed Daniel. I'm just looking at the scale of it. 100 sites. All from the same source. The amount of time involved to do that just doesn't seem to be a wise use of productivity to me at that scale.
I have to agree with Daniel's first suggestion, and at the same time, caution that the 2nd suggestion could rapidly backfire because just one link from one blog is very insignificant in the long run. Multiply that by multiple doctors - how long do you think you'd be able to get away with the blog comment path before someone smelled a rat and reported it as spam? In the blink of an eye.
Local citations, local directories, and local listings and Yelp, CitySearch and the rest. Coupled with high quality unique content on each site where you seed location info into the on-site optimization.
To add to Corey's response, I'll repeat what I just provided another question here on Pro Q&A. Sitemap.xml files can handle a maximum of 50,000 URLs, however I've seen them choke with as few as 10,000. Its important to run them through a tool like tools.pingdom.com to ensure they load within just a couple seconds.
Then submit them through Google/Bing webmaster systems and then see if they succeed in crawling all of them.
Dave,
Thanks for the clarification. You're definitely in a rare circumstance as compared to most web sites.
In reality, since it's the Bible, there is going to be a duplicate content issue regardless, given how many sites currently and how many more will most likely publish the same content now and in the future. From Eternalministries.org to KingJamesBibleOnline.org, concordance.biblebrowser.com, and so many other sites are all offering this content.
If you can find a way to offer your content in a unique way, and within your own site, offer different versions of it (individual verses compared to entire chapters), then ideally yes, you'd want it all indexed.
How you do that without adding your own unique text above or below each page's direct biblical content is the issue though.
Given this challenge,this is why I offered the concept of not indexing variations. Even if you weren't hit by the Panda update, any time Google has to evaluate multiple pages across sites where the content is either identical or "mostly" identical, someone's content is going to suffer to one degree or another. Any time it's a conflict within a single site, some versions are going to be given less ranking value than others.
So unfortunately it's not a simple, straight forward situation where duplication avoidance can be guaranteed to provide the maximum reach, nor is there a simple way to boost multiple versions in a way to guarantee that they'll all be found, let alone show up above "competitor" sites.
This is why I initially offered what are essentially SEO best practices for addressing duplicate content.
If you don't want to lose the traffic you have now that come in by multiple means, the only other way to bolster what you've got already is to focus on high quality long term link building, and social media.
The link building would need to focus on obtaining high quality links pointing to deep content. (Specific chapter pages and specific verse pages), where the anchor text used in those links varies between chapter or verse specific words, broader bible related phrases, and the LDS brand.
On the other hand, by implementing canonical tags, you will definitely reduce at least a number of visits that currently come in by variation URLs. Will that be compensated for by an equal or greater number of visits to the new "preferred" URL? In this rather unique situation there's no way to truly know. It is a risk.
Which brings me back to the concept that you'd potentially be better off finding ways to add truly unique content around the biblical entries. It's the only on-site method I can think of that would allow you to continue to have multiple paths indexed. Combined with unique page Titles, chapter/verse targeted links and social media, it could very well make the difference.
With what, over 1100 chapters, and 31,000 verses, that's a lot of footwork. Then again, it's a labor of love, and every journey is made up of thousands of steps. 
Dave,
You're facing a difficult challenge - satisfy the needs of SEO, or user experience. In light of all that Google has done going back to their May Day update last year and right through the Panda/Farmer update, duplicate content, as well as "thin" content, is more of a concern than ever.
Just having unique titles on each page is not enough. It's the entire weight of uniqueness.
Since you're not intending to go to individual pages for each verse, as long as you've got multiple methods of getting tocontent that is found by other methods, only one method should be designated as the primary search engine preferred method. All others should be blocked from being indexed.
From there, users can choose to explore other methods of finding content as they bookmark your site if they find it of help to their goals.
Unfortunately, this does of course, mean that you're going to end up with many less pages indexed. However every page that is indexed will become stronger in their individual rankings, and that in turn will boost all of the pages above them, and the entire site over time.
And here's another issue - when I go to any of the URLs you posted above, your site automatically tacks on "?lang=eng" using 301 Redirects. This means any inbound links you have pointing to the non-appended URLs are not providing maximum value to your site, since they point to pages designated as permanently moved.
Yeah I knew they were doing it with books, just wasn't aware about PDFs in general til Keri pointed it out.
Also - if the only change is the domain name, (URLs are otherwise identical, content is otherwise identical), then you may see more of a hit at Bing than Google simply because Bing emphasizes age of content a bit more.
Also I'd highly suggest a comprehensive link building initiative as the new domain goes live so while you're working to get as many existing links pointing directly (to bypass 301s on those 3rd party authority signals), you'll also want to have a variety of new quality links to show that "this new domain really is still authoritative and trustworthy". That should help ensure any hit you take is not as long as it would otherwise be.
I'd also encourage a press release sent out through PRWeb.com or PRNewswire.com, and social push announcing the change. All these are as much for signals confirming it's a legit transfer as they are for other SEO reasons.
Glad to be of help Janice.
From a readability perspective, in which case I'd suggest to have all lower case.
I would not recommend offering money for it. I would simply ask to be included, or at the very least, what the criteria are for being placed on the page. Links from such pages are not as valuable as you might think as far as SEO goes - they might be valuable if you believe a lot of people who are your ideal market go to such a page and click through.
Another factor is whether the page is clearly labeled as advertising space. If it is, then sure, a fee would be valid. But only based on statistics or data showing visitor usage and click averages, to justify the fee. If it's not clearly defined as ad space, paying violates Google's terms of service (and possibly the FTC's rules on disclosure) and should be done with extreme caution.
If the server allows upper case and lower case then from a technical perspective they could both be different files. Like having www.domain.com and domain.com point to the same home page - they may be the same, but technically they could be two different places.
The solution should be set up to not require having to do a rewrite every time a new page is created. It should be automatic.