I use a 3rd party tool that outputs a file in Unicode format. However, I prefer it to be in ASCII. The tool does not have settings to change the file format.
What is the best way to convert the entire file format using Python?
I use a 3rd party tool that outputs a file in Unicode format. However, I prefer it to be in ASCII. The tool does not have settings to change the file format.
What is the best way to convert the entire file format using Python?
It's important to note that there is no 'Unicode' file format. Unicode can be encoded to bytes in several different ways. Most commonly UTF-8 or UTF-16. You'll need to know which one your 3rd-party tool is outputting. Once you know that, converting between different encodings is pretty easy:
As noted in the other replies, you're probably going to want to supply an error handler to the encode method. Using 'replace' as the error handler is simple, but will mangle your text if it contains characters that cannot be represented in ASCII.
You can convert the file easily enough just using the
unicode
function, but you'll run into problems with Unicode characters without a straight ASCII equivalent.This blog recommends the
unicodedata
module, which seems to take care of roughly converting characters without direct corresponding ASCII values, e.g.is typically converted to
which is pretty wrong. However, using the
unicodedata
module, the result can be much closer to the original text: