Skip to main content
GameDev.net gamedev.net
🔒 Locked

Horizon Engine – C++20 3D FPS Game Engine with ECS and Modern Renderer

Started by bhdr26k Jan 11 at 10:29 AM 1 replies 1.7k views
Original Post
bhdr26k
bhdr26k


I’m working on an experimental 3D FPS game engine in C++20, aiming to deeply understand engine internals from first principles rather than just using existing frameworks.

Currently I'm strictly following LearnOpenGL docs.

This project focuses on: Entity-Component-System (ECS) architecture for high performance. OpenGL 4.1 rendering with a PBR pipeline, material system, HDR, SSAO, and shadow mapping. Modular systems: input, physics (Jolt), audio (miniaudio), assets, hot reload. A sample FPS game & debug editor built into the repo.

Repo: https://github.com/jackthepunished/horizon-engine

This isn’t intended to be a commercial rival to any commercial game engines.

it’s a learning and exploration project: understanding why certain engine decisions are made, and how to build low-level engine systems from scratch.

I’m especially looking for feedback on: Architecture choices (ECS design, render loop, module separation) Your thoughts on modern C++ engine patterns

What you’d build vs stub early in a homemade engine

Tips from experienced graphics/engine developers Criticism and suggestions are very welcome — it’s early days and meant to evolve. Thanks for checking it out!

RmbRT
RmbRT

Writing this as I browse through the project…

  • I would not use JSON for storing scenes. It's error-prone and slow (text parsing, memory fragmentation, dictionary walks, structure validation). I recommend this technique used by Valve: position-independent complex data structures that can be persisted (or sent over the network) without any serialisation or deserialisation. JSON is not human-readable contrary to popular belief, because anything even remotely non-trivial will be so large and deeply nested that you can't navigate it without a special JSON viewer. So if it's not trivially readable anyway, you might as well put the effort that you put into serialising and deserialising and validation into something that visualises the data in a more meaningful way (such as graphically or something). Using position-independent binary data is also way more compact than text could ever be.

    Using this paradigm everywhere also helps you make the game/engine multiplayer-ready very easily, as all data can just be sent and received without conversion, although depending on how you do it, you might still need a validation step (can be skipped by using 16-bit relative pointers and just placing things into a 64KiB sandbox).

    And for data structure versioning, I recommend using version namespaces and a method that takes an old data structure and turns it into one that is one revision newer. Then simply import the latest version namespace as the default.

  • Using memory arenas is great, but I would not recommend using STL containers. They have unreasonable requirements, leading to suboptimal performance. For example std::vector contains code for handling exceptions that might happen during a move constructor, and will roll back a resize operation in that case. Which requires that the old and new buffers exist simultaneously. If you use data-oriented design, where you just have plain old data everywhere instead of classes with constructors and destructors, it would be better to invest into writing containers that are realloc()-aware, and operate on raw data.

    Another paradigm I would recommend, if you absolutely need something like a move constructor, is to have a function that instead of move-constructing from an existing copy, simply takes a raw-copied instance of an object, and tells it the address it got moved from. This is realloc()-friendly:

    T * new_objects = realloc(objects, new_size * sizeof(T), alignof(T));
    if(new_objects != objects)
    	for(unsigned i = 0; i < min(size, new_size); i++)
    		// objects[i] no longer exists! This only uses the address, not the memory it points to.
    		new_objects[i].was_relocated_from(&std::cref(objects[i]));
    objects = new_objects;
    size = new_size;

    Instead of move-constructing the new data in all cases (which requires source and destination to exist simultaneously), in case the address did not change in the reallocation, you can do nothing. And you can use memcpy() to copy objects, which is faster than memberwise copying. Additionally, you do not need to call the destructors on the old data, as it simply doesn't exist anymore. All you do is adjust the values in the copy that need to adapt to the changed position of the data. But for position-independent data types as recommended above, you would not even need to use this step. Having data that is trivially movable lets you have much more efficient and leaner containers.

  • In the ECS system, I would do the “sparse” lookup differently: First of all, I would add a method that lets you look up an entire batch of entities. You can look up at least 8 items at once without additional cost in runtime, as most computers can handle 8 or more simultaneous cache misses, afaik (# of line fill buffers / LFB). This would give you an 8x speedup for looking up multiple entities at once. Additionally, I would use a compression scheme for the sparse array that uses one bit per entity, instead of 32, and then uses the following structure:

    struct Entity_SparseEntry // Represents a chunk of up to 32 entities.
    {
    	alignas(8) // make sure entries do not cross cacheline boundaries.
    	uint32_t is_present = 0; // bool[32]
    	uint32_t base_index = 0; // index in the pool of the first entry of this chunk.
    	
    	// sub_index: [0, 31]
    	uint32_t get_index(uint32_t sub_index) const {
    		// feel free to make this branchless...
    		if(is_present & (1 << sub_index))
    			// count the 1-bits before the entity: this is our offset into the chunk.
    			return base_index + std::popcount(is_present & ((1 << sub_index) -1));
    		else return INVALID_ENTITY;
    	}
    };
    struct SparseEntityComponent {
    	Entity_SparseEntry sparse_map[];
    	Entity entities[];
    	ComponentArray components;
    
    	bool contains(Entity entity) const { return sparse_map[entity / 32].is_present & (1<<(entity % 32)); }
    	typename ComponentArray::ElemRef lookup_component(Entity entity) {
    		uint32_t idx = sparse_map[entity / 32].get_index(entity % 32);
    		if(idx != INVALID_ENTITY)
    			return components.get_elem_ref(idx);
    		return nullptr;
    	}
    };

    This is 16× more compact but needs a more sophisticated strategy for putting entities into the array, as you have fixed bucket groups, but a variable number of entities corresponding to that group.

    Additionally, I don't think it is good to have an array of structs for your entity components. A struct of arrays would be much more performant, as it allows you to use SIMD to process entities in bulk. The most expensive part of your game logic is always bulk operations that affect all entities, not one-off operations that operate only on one entity. For that, I recommend injecting a ComponentArray, not Component type, and having that one manage the various sub-arrays.

    Then, when you want to do bulk operations, you can loop over the component array and perform SIMD operations on it. I recommend having one such sparse entity component set per filter condition that you have in some bulk operation. Some bulk operations then operate on multiple such sets. The ability to use SIMD in those loops is definitely worth it. As they say: “Don't build a rocket class, build a rocket manager.” And some sets contain multiple components per entity, such as a set containing a position array and a velocity array for things that move, while another component set would only contain the positions of static entities. If then something wants to loop over all entity positions, it would have to loop over both sets, but only look at the position component in both sets.

    Also, ideally, your entity IDs could be namespaced, instead of everything sharing one polymorphic handle. That way, you can cut down the sparse map capacities by a lot. Let the game decide on the taxonomy of entities, as well as the number of entity sets. You can also have one additional array for all instances of a component that contains the ID of the component set in which it is currently located (uint8 should suffice, probably). This way, you have highly efficient SIMD bulk operations, and also efficient per-set lookups and efficient sparse storage, but also generic per-Entity lookup of components, regardless of in which bulk operation specific component set they are located. And with the new VPGATHER instructions, you could also use SIMD for scattered lookups.

    Also, not all component sets need a by-ID lookup. You should offer one version that has a sparse lookup and one that does not. The game should not pay a price in performance or resources for things it does not use.

  • In general, for an engine to be useful, it needs to do the job it is needed for really well. You should probably set a specific goal of what you want to achieve with the engine. Just bolting lots of features onto it will probably make those features be of lower quality overall than what flagship engines offer, without excelling at anything specific. The best way to make a good engine is to make the engine for a specific game only. Additionally, I would make the game handle as much game-specific stuff as possible, and not force it to use the entity system, for example, but rather offer it as a tool for the game. Because the more specific knowledge you have, the more optimised code you can write. And you as the generic engine developer don't know all the specific constraints and guarantees that the game has, and maybe some of these constraints allow some work that your generic solution has to do, to be skipped entirely. The best optimisation is doing less work, but that is only possible by understanding the specifics of the problem at hand. Which you by definition can't.

I'm also currently writing an engine, and I'm targeting OpenGL ES 2.0 upwards (specifically WebGL compatibility) with an abstraction layer that maps to modern OpenGL if available, but also cleanly maps onto old OpenGL, so that it runs on phones, and in the browser, and on desktop and consoles. The engine handles audio, graphics, remappable input, networking, and asset management, and runs in webassembly or native. This way, I can ship browser builds for playtesting without people having to trust the software they run, and it also gives me free portability across basically all platforms, so I don't have as much pressure to add native support for many platforms right away. Later on, I also plan to build a very small linux distribution that is basically just my engine, and nothing else (except also my compiler). Anyway, my engine is separated into a general runtime layer that interfaces with the hardware, and then a game-engine specific layer that handles game-specific functionality and data structures. I push some responsibilities of the engine into the game, because the game has more specific knowledge and can make more informed decisions (such as choosing the right vertex formats for models or the appropriate compression schemes for textures, etc.). Basically the game is responsible for using the functionality provided by the engine to build another, highly optimised and specific engine on top of that. The job of the outer engine is to make the building of that second engine as easy as possible. Building a game directly on top of a generic engine that is not designed to be a secondary engine introduces too much friction because the abstractions & interface usually don't match with the requirements of the actual game.

Edit: Sorry if there are any corrupted sections, I just noticed that it had somehow deleted parts of my code sample. I rewrote that part but it might also have deleted parts of the rest of the post. This site sometimes drives me nuts.

Walk with God.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.