Skip to main content
GameDev.net gamedev.net
🔒 Locked

How do I save an std::string to a file if there are non-ascii characters present

Started by Drakkcon Sep 3, 2009 at 5:24 PM 6 replies 3.2k views
Original Post
Drakkcon
Drakkcon
I'm using std::string in my project, and I have selected "Unicode" as my character set in visual studio, which I assume means UTF-8. Now, I want to just be able to do this:
stream.write(string.data(), string.size());
But looking at documentation about string has made me uneasy, since every reference I have found states that size() is the same as length() and gives you the number of characters in the string. However, in the case of most strings containing non-ascii characters, the size of the buffer is different from the number of characters. How do I get the size of the actual buffer? If there is no way, how does one save a unicode string to a file in c++?
Sc4Freak
Sc4Freak
No, that's not what the "Unicode" setting does. The "Unicode" setting in VS just #defines UNICODE project-wide. This can affect things like WinAPI calls, so that DrawText for example will be routed to DrawTextW (Unicode) rather than DrawTextA (ANSI).

std::string and std::wstring don't necessarily correspond to ANSI and Unicode strings, respectively. It's just that std::string uses a narrow characters (char) and std::wstring uses a wide char (wchar_t). On Windows, Unicode is represented using UTF-16, which is why wchar_t is 2 bytes. But since std::wstring isn't specifically UTF-16, size() and length() don't take into account surrogate pairs and the like.

The short of it is: Unicode in C++ is a mess. size() and length() will get you the number of code units in the string, not the number of code points. This is because neither std::string nor std::wstring know which encoding is being used.
stonemetal
stonemetal
I typically go with the old stand by of ofstream << string, but maybe that is just me.

Quote:
size() is the same as length()
Yes.
Quote:
gives you the number of characters in the string.
only if you use ascii. I am not sure but I think Unicode on windows usually means UTF-16. std::string is actually a template specialized on chars so length\size return number of chars and has nothing to do with characters. This also means that you could specialize it on something else to get size\length to represent something else. Mostly std::string is vector with a pretty face.
Drakkcon
Drakkcon
Quote:

The short of it is: Unicode in C++ is a mess. size() and length() will get you the number of code units in the string, not the number of code points. This is because neither std::string nor std::wstring know which encoding is being used.

Oh, so size() and length() are the length of the buffer? Awesome, that's what I was hoping (although I can see how it would suck if you actually wanted to iterate over each character).

Quote:
I typically go with the old stand by of ofstream << string, but maybe that is just me.

Does this work with binary streams? If so I will use it.

Thanks a lot both of you!
Zakwayda
Zakwayda
Quote:
Oh, so size() and length() are the length of the buffer? Awesome, that's what I was hoping (although I can see how it would suck if you actually wanted to iterate over each character).
size() and length() return the number of characters in the string, not the size of the buffer (of course in some circumstances the number of characters and the number of bytes in the buffer will be the same).
Codeka
Codeka
Quote:
Original post by jyk
Quote:
Oh, so size() and length() are the length of the buffer? Awesome, that's what I was hoping (although I can see how it would suck if you actually wanted to iterate over each character).
size() and length() return the number of characters in the string, not the size of the buffer (of course in some circumstances the number of characters and the number of bytes in the buffer will be the same).
Number of code units, not characters. For example, "é" is one "character", but two "code units" (U+0065 U+0301). It can also be written as "́é" which is one code unit (U+00E9). Welcome to the wonderful world of Unicode :-)
Drakkcon
Drakkcon
Alright cool, and each code unit in UTF-8 is a byte so I can be like:

stream.write(str.data(), str.size());

and it'll work right. Thanks everyone, for your help.


Edit: and by work right I mean it will copy the entire buffer through the stream without clipping it off if there are multi-byte characters.
Codeka
Codeka
Quote:
Original post by Drakkcon
Alright cool, and each code unit in UTF-8 is a byte so I can be like:

stream.write(str.data(), str.size());

and it'll work right. Thanks everyone, for your help.

Edit: and by work right I mean it will copy the entire buffer through the stream without clipping it off if there are multi-byte characters.
Correct.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.