Criticizing OSM is at least as hazardous as criticizing Wikipedia, but...
Setting aside ideological purity, the actual rationality of switching solely to OSM depends on the actual quality and freshness of its data. And most of the world with money in play is more likely to follow rationality than ideological purity here.
But with regard to data quality, OSM is encumbered its own policy. OSM cannot (or refuses to) make use of high-quality bulk data from outside, regardless of how that data was licensed. Automation allows more frequent updates with a lot less manual labor as input - but OSM requires contributors to manually trace shapes in their weird browser app. Even where plenty of excellent liberally licensed data is available, OSM insists on making an army of unpaid monkeys re-make the SAME data in perpetuity.
That puts its pipeline at a permanent disadvantage to companies building proprietary datasets, which can take advantage of all kinds of automation. OSM simply isn't doing as much as it could to verify its data and stay fresh, and so it isn't doing as much as it could to end the status quo of proprietary data. Even if OSM reaches a usable state at a given moment, the world keeps changing to invalidate previously traced data, and OSM's policy ensures it will always be playing catch-up no matter who wants to donate data, because all the data has to come through this dopey manual process.
This viewpoint seems to be based on the assumption that proprietary datasets are going to be more up to date than a crowdsourced dataset. Having worked with proprietary datasets, often this assumption is wrong as these datasets (which could be eligible for importation) are often surprisingly poor quality and out of date. If a company can still make money selling their datasets to their same customers, there is no incentive for them to spend resources to keep it up to date.
What this viewpoint does hint at is it's comparison to Wikipedia, and the differences in geospatial data of a current map of the world. Pekk is correct, an area in OSM needs to be continually updated and refreshed, and this requires appropriate tools and a user base. Both of these exist or are being improved upon with OSM.
For example, work has happened re: Monitoring of changes, comparisons with imagery, map bug reports, automated feeds, academic and industry reports and evaluations of the quality with proprietary and national mapping agency datasets.
This is very true. My country's Government had an issue with a major internet project where they bought the best commercial address database available. Turns out that hundreds of thousands of addresses in there were incorrect, and tens of thousands just didn't exist. Caused huge problems.
Proprietary geospatial data can be very hit and miss. If you assume that the majority will be high quality, you're setting yourself up for disappointment...
> * OSM cannot (or refuses to) make use of high-quality bulk data from outside, regardless of how that data was licensed. *
There is actually not a lot of "high quality bulk data" out there. What datasets are you thinking of?
Remember OSM is "open data", so you can't import propriatary data. Some bulk datasets might not allow commerical reuse, or might forbid derivative works (i.e. no-one can change it).
> OSM requires contributors to manually trace shapes in their weird browser app
There are many editors. iD is the browser app. There's JOSM a desktop, advanced app. And you can always upload raw data via the API.
OSM is constantly importing bulk data sets, anyway. For instance, in the UK the NAPTAN database was imported, giving the location and reference codes for all public transport stops. Postcode (like ZIP code) data is imported from a monthly building sales data release. I noticed the other day that Northern France suddenly has all buildings marked, because of some government data import. There are bulk data imports happening every week somewhere in the world.
Just for the record - NAPTAN wasn't consistently imported in the UK; we tried for a few counties and found it wasn't that great. The UK is probably the most import-hostile national community in OSM (and also, after Germany, probably the most successful in building the map; up to you whether you see a coincidence there!).
Is the purpose to make good data, or to make a community of people manually tracing things that were already traced in liberally licensed data? And how exactly does this help with maintenance?
It has been found that encouraging a community around the data collection creates much better data in the long run. An involved community can go beyond what is provided by other datasets and do more than can be done with 'tracing'.
There's more to OSM that road centre-lines and building outlines.
Can you link to any examples of them tracing over bulk, good quality, correctly licensed data? For example wiki pages or mailing list discussions saying they require this?
I think the effort is to make a community. You can always make your own maps separately that mix in other data. Trying to make a single map that will be all things for all people is a waste of time. Streetmaps only every show a tiny amount of data that could be shown.
Why do you think OSM isn't capable of utilizing automated processes? If anything, the manual editing process allows OSM to catch the long-tail data touch-ups that automated processes will never be able to handle.
Nobody said there should never be manual touch-ups, but you can certainly do manual touch-ups on top of bulk imports. Everyone else does. A process which depends on a ton of unpaid manual labor to do what bulk imports should be doing is not very scalable... if you have that labor, why not ensure that it goes to the touchups rather than to duplicating things people have already done and want to contribute
>A process which depends on a ton of unpaid manual labor to do what bulk imports should be doing is not very scalable...
Oh. I think yes this is scalable. I'll give folks some examples of how crowd-sourcing works like this, and how it can be scalable.
1. A person knows where they live, they can add their own street. They may know their neighbourhood, they can map it from scratch, or more likely keep it up to date. The user may see that their family home or where they vacation needs a tweak here and there, and pops on and edits. It's very casual, very local based monitoring and editing. This is scalable in the sense that it's small tasks that take a small time.
2. Based on an academic evaluation of a countries government's base map and OSM, several places are shown to be less well covered in OSM. These are targets for more dedicated mappers to examine, on foot, or with aerial imagery to improve. Hardcore mappers mainly here, large tasks, large time.
3. Mapping workshops are held in a small town to blitz a new development, fill holes and cover an entire area all at once. Twenty participants can do a large area with transport and the right equipment.
4. web based task manager is set up to allocate squares of land for people to trace from, dozens of people can take a small square and trace over the top. Local knowledge can then quickly add in names and landmarks.
So you can see that there are a number of different crowd-sourced labour approaches which would replicate what an import would do, but that by crowdsourcing you actually increase the number of people in that labour pool. An import does not help increase the number of contributors.
Now, this is more scalable in developed countries where there are more people with time to donate (same applies to any unpaid crowdsourced project), and so that example no. 1 in developing countries of an organic community keeping things up to date wouldn't really happen. However it's quite widespread in the developing world that there are no proprietary datasets of any decent quality. Imports could not even be considered in these places.
The problem with manual touch-ups on imported data is that when a newer version of the bulk data is released there's a difficulty with how to merge the datasets. This is the problem the US mappers are finding with the newer versions of TIGER.
And there are in fact huge amounts of bulk-imported data despite the wariness. Much of the U.S. road network is based on TIGER imports (especially outside cities that have a lot of OSM editors). Some of the bulk imports are easy to see on a macro scale, because there will be a conspicuous change in data density at a political boundary: http://www.openstreetmap.org/#map=8/34.266/-84.468
I think it varies by language and country too. The UK had to start from zero (and this prompted the creation of OSM) but may have left a cultural impact on the kind of person who joined the project, whereas the US had freely licensed Tiger data that was really created for an entirely different purpose and created a lot of problems and stopped a community forming (or at least that's the common storyline, I think it's more complex than that).
Whereas some other countries seem to have had better sources of data available and got on with the job of figuring out the best ways to import them without making a mess.
Yeah, it's a tricky question, especially when it comes to updating (imo the initial import was fine, though some people disagree).
I like the current assisted approach of letting you turn on a TIGER 2012 overlay in the editor, and/or use tools like the "TIGER 2012 Battlegrid" (http://maproulette.org/battlegrid/) to find where OSM/TIGER differ.
So far based on some limited experience fixing OSM based on those tools, it's a good decision imo not to do it in an automated fashion, because when I've found OSM/TIGER differences, it was about 50/50 which of the two was right.
And again, OSM is not fighting for the biggest user-share, but for the freedom of its users - who are not supposed to just be users, but maintainers too. It is a project for freeing all the human people from any lock-in for all geographical data.
See also the recent buzz around gcc (TL;DR: gcc is about freeing users from proprietary implementations, not about having the best product ever), and Wikipedia (everybody knows something, Wikipedia is here to share this knowledge and make sure it can't be withheld)
You seem to have defined rationality to mean 'a person who cares solely about the present, with no value put on the future'.
Because a person who rationally cares about great map data for their whole lives, and also cares about other things (like liberty) might well prefer the growing pains of being an earlier adopter to the instant convenience of using Google Maps.
I understand that you, like many people, have decided that the attributes you care about are the only ones that are rational to care about. But that is a gross mistake, and that mistake undermines the entire premise of your rant.
This isn't about my preferences. The whole point is that only an arrogant asshole would claim their preferences are the only rational ones.
As another example within GPS, Waze used to give truly horrendous routes, as it would often fail to understand the difference between overpasses and intersections. A bunch of early adopters liked the promise of Waze and helped crowd-source the fixes, so that now Waze is a very good navigator.
By the OP's logic, the early adopters who helped make Waze great were irrational, because it truly was worse for quite a long time.
That said, I get it, you're just here to snark at strangers. After all, if you'd had even a hint of an interest in a real conversation you would've realized I never stated that any set of attributes was better than any other, nor that my interpretation of attribute definitions was the only valid one.
Anyway, thanks for the snarky bullshit and have a nice day.
Setting aside ideological purity, the actual rationality of switching solely to OSM depends on the actual quality and freshness of its data. And most of the world with money in play is more likely to follow rationality than ideological purity here.
But with regard to data quality, OSM is encumbered its own policy. OSM cannot (or refuses to) make use of high-quality bulk data from outside, regardless of how that data was licensed. Automation allows more frequent updates with a lot less manual labor as input - but OSM requires contributors to manually trace shapes in their weird browser app. Even where plenty of excellent liberally licensed data is available, OSM insists on making an army of unpaid monkeys re-make the SAME data in perpetuity.
That puts its pipeline at a permanent disadvantage to companies building proprietary datasets, which can take advantage of all kinds of automation. OSM simply isn't doing as much as it could to verify its data and stay fresh, and so it isn't doing as much as it could to end the status quo of proprietary data. Even if OSM reaches a usable state at a given moment, the world keeps changing to invalidate previously traced data, and OSM's policy ensures it will always be playing catch-up no matter who wants to donate data, because all the data has to come through this dopey manual process.