Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"In 2010, Google estimated that it had indexed just 0.004% of the internet."

I don't believe this. Does anyone else?



This sentence made me bug too. One should not mix up the web and the internet. Google only indexes the web (well, and a part of usenet). Found this: http://webapps.stackexchange.com/questions/11740/how-much-of... which support the claim in the article.


It's probably because Google mostly indexes public-accessible text and images. I have a home server with a few TB of Internet-accessible storage, but it's not public. I can say the data is on the Internet, but cannot be indexed as it's not directly accessible. My cloud backup of that storage is on the Internet but it's not public and it's encrypted, so it cannot be indexed (maybe the NSA is crunching my 4096-bit key right now). Based on the infographic agilebyte provided (sibling post), I contribute to ~10TB of data that cannot be indexed by Google - and that's already 5% of the data Google had indexed by 2010!


There are some giant internet databases behind a "Disallow: /".


Infographic with source at the bottom here: http://urlm.co/blog/2010/10/28/the-awesome-size-of-the-inter...


So Google has indexed 200TB? My cat pictures alone are 1TB!


A shockingly high percentage of the internet is spam or other junk that is never seen. I don't know about .004%, but I don't doubt that Google doesn't index the absolutely huge volumes of crap out there.


Google may index a lot of the spam as relatively useless since it is often found at top levels. If the number is correct what it is probably missing is all of the content that lies in databases, easily accessible, but not necessarily on the front pages of web sites, or found without entering a query of some sort.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: