The difference is between raw performance and "voicing".
You can say "this model of headphones produces peaks at 4kHz, 11kHz and 26kHz compared to a flat reference model" and add that effect with a DSP or pre-rendered sound clips.
That will give you some of the "voice" but it won't add the performance aspects-- power handling capacity, responsiveness of the drivers, sharpness of the crossovers, etc.
TBH, it's a fascinating idea and I wish it had existed last time I bought speakers-- I recall going all over town to Best Buys and Fry's back when they still had stock to listen to floor samples in rooms completely unlike the ones I would be listening in, using sample media and amplification unlike what I was going to use.
If it could then you would just use the software on all your music and have high end audio for free.