Seriously, nobody would do that. RMXP doesn't even have the possibility to handle audio input, and I'm not sure if RGSS does, so it probably would need external coding. Then, somebody would need to "teach" the program how to handle the imput and the recognition engine, which is pretty much impossible to do for probably nearly all RMXP coders. You can't just directly compare the Audio imput to a file in the database, since everyone's voice is different, but the program needs to analyze the relativr pitch and volume changes, recognise and ignore unwanted noise, and much more. Big companies spend years to get a decent voice recognition engine, and you expect a single programmer to do it in mich less time, for free, in an engine that isn't even worth it?
tl;dr Forget about it.