Skip to main content
GameDev.net gamedev.net
🔒 Locked

Console Advice

Started by lordimmortal2 Feb 3, 2011 at 12:51 AM 6 replies 1.4k views
Original Post
lordimmortal2
lordimmortal2
Over the past week, I have been trying to build a developer console to learn just how to do it (and I didn't know if there was any other method than to build one in-house). But, as I got farther and farther into it, it got more convoluted and confusing and I'm starting to wonder if there's a better way to do it.

So, what I want to ask is: is there any DirectX, Windows or something similar like a function that will be a good way to manipulate text? Or is there a tutorial on how to build a developer console?

Basically is there any better way than building the developer console in-house or is that the only way?
Prove me wrong so I can know what's right.
SiCrane
SiCrane
For dealing with the text manipulation I suggest choosing a scripting language and letting it worry about string processing. I personally prefer python, which in particular has some nice interactive console support.
Ravyne
Ravyne

Over the past week, I have been trying to build a developer console to learn just how to do it (and I didn't know if there was any other method than to build one in-house). But, as I got farther and farther into it, it got more convoluted and confusing and I'm starting to wonder if there's a better way to do it.

So, what I want to ask is: is there any DirectX, Windows or something similar like a function that will be a good way to manipulate text? Or is there a tutorial on how to build a developer console?

Basically is there any better way than building the developer console in-house or is that the only way?


There really shouldn't be anything that difficult about it.

You'll need:
  • A way of getting full keyboard input.
  • A workable solution for text rendering (there are many).
  • A means of tokenizing the input strings.
  • A means of parsing the token stream.
  • a means of executing the parsed commands.
    The tokenizer is fairly straight-forward, basically you scan characters left-to-right, taking as many characters into a "token" as are in the same category (letters for words, digits form numbers, strings are anything between two quote marks, etc) -- then you record the type of the token (word, number string) paired with its string value (repeat, 4, "Hello, World!"). Regular expressions are extremely handy here, even for grammars of a fairly complex nature... Regex should be capable of tokenizing any grammar where the the type of a token cannot be affected by preceding tokens.

    The parser is the most difficult, but I would say that most interactive console commands can be parsed with a simple recursive-decent parser. Basically recursive-decent parser looks at one token at a time, in order, and determines the next set of possible tokens based on that -- if the next token isn't one of the expected ones, its an error. The way you'd probably structure this in a console is defining the set of global command tokens (things your console can 'do', sometimes you call these verbs in grammar parlance), and each of these has its own parse function which is handed the token stream and consumes as many tokens as it needs, according to its own rules -- in other words, the "repeat" command from above knows that it should consume a number, and then a string. If it finds them as expected in the next two tokens (order counts) then it takes them from the token stream and has all the information it needs to execute the command (usually, each command has a native-code function to perform its action) so you do any last necessary conversions (string of a number to a real number, for example) and call that function and return. You might also need to pass in a console context implicitly to each of these native code handler functions so that it can print to the console, or perhaps read state (variables) that have been set in the console. If at any point a command finds a token it doesn't know how to deal with, then there's a syntax error in the command.

    There are systems for generating tokenizers and parsers, however they are often non-trivial to learn, and most of them spit out vanilla C code littered with gotos and non-sensical labels and variables. For large or complex grammars, the learning curve can be a worthwhile investment. But I don't think there's much value in them for anything that can be handled by a simple recursive-decent parser -- its probably just as quick to build one by hand than to use the tools, so unless you plan to make a habit of using the tool I wouldn't recommend learning it.

    If you do deem it of sufficient necessity (or curiosity) then the typical tools are lex/yacc, flex/bison, Boost::Spirit and ANTLR among many others.
throw table_exception("(? ???)? ? ???");
lordimmortal2
lordimmortal2

[quote name='lordimmortal2' timestamp='1296694283' post='4768804']
Over the past week, I have been trying to build a developer console to learn just how to do it (and I didn't know if there was any other method than to build one in-house). But, as I got farther and farther into it, it got more convoluted and confusing and I'm starting to wonder if there's a better way to do it.

So, what I want to ask is: is there any DirectX, Windows or something similar like a function that will be a good way to manipulate text? Or is there a tutorial on how to build a developer console?

Basically is there any better way than building the developer console in-house or is that the only way?


There really shouldn't be anything that difficult about it.

You'll need:
  • A way of getting full keyboard input.
  • A workable solution for text rendering (there are many).
  • A means of tokenizing the input strings.
  • A means of parsing the token stream.
  • a means of executing the parsed commands.
    The tokenizer is fairly straight-forward, basically you scan characters left-to-right, taking as many characters into a "token" as are in the same category (letters for words, digits form numbers, strings are anything between two quote marks, etc) -- then you record the type of the token (word, number string) paired with its string value (repeat, 4, "Hello, World!"). Regular expressions are extremely handy here, even for grammars of a fairly complex nature... Regex should be capable of tokenizing any grammar where the the type of a token cannot be affected by preceding tokens.

    The parser is the most difficult, but I would say that most interactive console commands can be parsed with a simple recursive-decent parser. Basically recursive-decent parser looks at one token at a time, in order, and determines the next set of possible tokens based on that -- if the next token isn't one of the expected ones, its an error. The way you'd probably structure this in a console is defining the set of global command tokens (things your console can 'do', sometimes you call these verbs in grammar parlance), and each of these has its own parse function which is handed the token stream and consumes as many tokens as it needs, according to its own rules -- in other words, the "repeat" command from above knows that it should consume a number, and then a string. If it finds them as expected in the next two tokens (order counts) then it takes them from the token stream and has all the information it needs to execute the command (usually, each command has a native-code function to perform its action) so you do any last necessary conversions (string of a number to a real number, for example) and call that function and return. You might also need to pass in a console context implicitly to each of these native code handler functions so that it can print to the console, or perhaps read state (variables) that have been set in the console. If at any point a command finds a token it doesn't know how to deal with, then there's a syntax error in the command.

    There are systems for generating tokenizers and parsers, however they are often non-trivial to learn, and most of them spit out vanilla C code littered with gotos and non-sensical labels and variables. For large or complex grammars, the learning curve can be a worthwhile investment. But I don't think there's much value in them for anything that can be handled by a simple recursive-decent parser -- its probably just as quick to build one by hand than to use the tools, so unless you plan to make a habit of using the tool I wouldn't recommend learning it.

    If you do deem it of sufficient necessity (or curiosity) then the typical tools are lex/yacc, flex/bison, Boost::Spirit and ANTLR among many others.
    [/quote]

    Thank you for your reply, this is very very useful information. Normally when I try to do something by myself, it's full of convolutions, strange integers, and gets really confusing. I don't think I could ever think of something like this off the top of my head.

    Would just using DirectInput's KEY_DOWN function that fills a data structure for all the keys of the keyboard be sufficient for the full keyboard input or is there a more efficient way?

    I already have a method for rendering the text.

    Is there a sort of tutorial around for how to build a tokenizer? I get it a fair bit, but I'd probably mess it up trying to build it. What specifically is a token? An array of characters? And for the parser too? I get the parser more than the tokenizer though.
Prove me wrong so I can know what's right.
KulSeran
KulSeran

For dealing with the text manipulation I suggest choosing a scripting language and letting it worry about string processing. I personally prefer python, which in particular has some nice interactive console support.

I'm going to second this.
Just worry about taking key inputs and building a string. Once you get an "enter" key press, pass the string off to your lua / python / anglescript interpreter. Let the scripting language deal with all the tokeninzing and parsing and binding of variables to function calls.


Is there a sort of tutorial around for how to build a tokenizer? I get it a fair bit, but I'd probably mess it up trying to build it. What specifically is a token? An array of characters? And for the parser too? I get the parser more than the tokenizer though.
[/quote]
You can use tools like flex + bison, or boost::spirit to deal with most of this stuff for you.

It involves runing the input string through a set of regular expressions that represent valid forms of "token". So, you'd have [0-9]* is an integer, while [0-9]*\.[0-9]*[f] would be a float and [a-zA-Z_][a-zA-Z_0-9] would be a variable/function/item name. You then send the tokens into a parser that can turn rules like "function_call := identifier paren expression_list closeparen" into your abstract syntax tree that you can then walk to find your function name and parameters.

For a console though, you can usually write much more simplistic parsers and tokenizers by limiting the types of statements you can enter at the console. It's still a lot more trouble than just sending it all off to a scripting language. But you could take a look here on gamedev. And google for "cvar console" or "quake console" as those are the more common term for a game's developer console.
Ravyne
Ravyne

Would just using DirectInput's KEY_DOWN function that fills a data structure for all the keys of the keyboard be sufficient for the full keyboard input or is there a more efficient way?

I already have a method for rendering the text.


I may be wrong, but I believe that these types of functions usually are just giving you a state that represents whether a key was pressed or released since the keyboard was last poled -- or, perhaps it does work off of keydown/keyup type events, but that this information is lost. For game input, you normally don't care so much about the order in which keys were pressed or released, as long as we're within a single frame. Obviously, for typing text the order of presses and releases is very important. I imagine you'll need to hook into the key messages as they occur. Direct3D may provide this functionality in some form, but I don't think so. Other frameworks (SDL, SFML, etc) might provide it... I'm not familiar with either. Worst-case scenario, you should be able to tap into the Windows Event system to process the messages in order.


Is there a sort of tutorial around for how to build a tokenizer? I get it a fair bit, but I'd probably mess it up trying to build it. What specifically is a token? An array of characters? And for the parser too? I get the parser more than the tokenizer though.


Any book on writing text processing, scripting languages or compilers will treat tokenizers extensively. A Google search for Tokenizer/Tokenization should turn up good info. Also, Tokenization almost always shows up as a sort of preamble to any resource on parsers or parsing -- mostly because tokenization is a necessary step to parsing, but doesn't do much good on its own.

Tokens themselves are very simple creatures. At the bare minimum a token is simply a string which contains text from a larger string. The token string was simply singled out from the larger string as being some kind of cohesive "thing" (a word, a number, a quoted string, whitespace). Usually though, you want to take advantage of the fact that you've figured out what category of "things" it belonged to already, so you store that along with the string so you don't have to figure it out again later; we call this the token's "type". So there you have it, a token is just an object that contains a type and a string. Sometimes people will optimize for space by only including the string only if the string cannot be inferred directly from the token (eg, if you have a token who's type is 'COMMA', then you can simply assume that the string is ",", rather than storing it every time you parse a comma.), but this is an optimization. Get things working first before you worry about a few lost bytes here and there.

Lets assume we've defined a tokenizer that recognizes words, numbers, quoted strings, and whitespace. As part of that system assume we've included std::string, and that we've defined the following types:



enum token_type_t {
EOF, // End of File
WHITESPACE, // Space, Tab, Return/Newline
WORD, // any contiguous characters that are not EOF, WHITESPACE, NUMBER or inside a STRING
NUMBER, // any contiguous digits
STRING // any characters, digits or whitespace between quotation mark pairs
};

struct token {
token_type_t type;
std::string text;
};



If we turn our tokenizer loose on the string ["Holy Cow!" That's more than 100!"], we would get a list of tokens back like this:






">

Finally, if you're willing to invest in a book, I've been reading a book called "Language Implementation Patterns" which covers several different techniques for tokenization, parsing and beyond (the book itself addresses creating scripting and programming languages, so it then goes into managing scopes, typing, intermediate representations, analysis, optimization and code generation) -- its more than you need for a simple developers console, but it is quite exhaustive of the parts you do need and is also quite approachable. If you were instead to pick up a book on creating a compiler, you would find that they soon become lost in a world of grammars, state-machines, graphs, etc... Useful stuff, but stuff that would get in the way of a relatively simple goal in your case.
throw table_exception("(? ???)? ? ???");
SiCrane
SiCrane

Would just using DirectInput's KEY_DOWN function that fills a data structure for all the keys of the keyboard be sufficient for the full keyboard input or is there a more efficient way?

Just use the standard WM_CHAR window message. It'll be a lot easier than using DirectInput.
lordimmortal2
lordimmortal2
Thanks everyone for your replies. I'll try each suggestion and see if I can't get a console working. If I run into any problems, I'll be sure to post here again. Thanks again =).
Prove me wrong so I can know what's right.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.