Skip to main content
GameDev.net gamedev.net
🔒 Locked

Why did no one tell me about how useful the is into making games?

Started by St3f0n May 10 at 1:47 PM 19 replies 2.1k views
Original Post
St3f0n
St3f0n

Hey everyone,

I've been experimenting a lot with using AI in my game dev workflow — specifically crafting prompts that actually produce useful results for game design, worldbuilding, art direction, and more. After a lot of trial and error I put together a collection on my website, but wanted to share a few examples here first.

─────────────────────────
Text prompts (ChatGPT / Claude)
─────────────────────────

🎮 Enemy AI behavior:
"Design a boss enemy for a dark fantasy RPG. Give it 3 phases, each with a distinct attack pattern and a visual cue that telegraphs the next phase. Include a weakness the player can discover through exploration."

🗺️ Level design brief:
"Generate a level design brief for a stealth game set in a medieval castle. Include entry points, patrol routes, 3 optional objectives, and one environmental hazard the player can use to their advantage."

📖 Lore & worldbuilding:
"Write a 200-word in-world legend about the origin of a cursed artifact. The tone should feel like it was passed down orally — imprecise, slightly contradictory, and ominous."

─────────────────────────
Image gen prompts (Midjourney / Stable Diffusion)
─────────────────────────

🎨 Environment concept:
"A foggy swamp village built on wooden stilts, dark fantasy, muted greens and browns, volumetric lighting, concept art style, highly detailed"

⚔️ Character design:
"A female mercenary in worn leather armor, scarred face, holding a crossbow, neutral pose on white background, character sheet style, 2D game art"

🏚️ Prop / item art:
"An ancient spell tome with cracked leather cover and glowing rune clasps, flat lay view, isolated on white, stylized hand-painted game asset"

─────────────────────────

There are a lot more on my site covering mechanics brainstorming, dialogue writing, UI mockup prompts, and more: https://devloot.gumroad.com/l/gamedevaipromptpack

Would love to hear what AI tools people are using in their pipelines — and if you try any of these, let me know how they turn out!

Cheers

JoeJ
JoeJ

St3f0n wrote:

Would love to hear what AI tools people are using

I don't use artificial intelligence. I'm intelligent myself, so i don't need it.

Same about 'generative AI' regarding creativity.
In case i would not have any ideas about lore, levels, or bosses for my game, i would just make another game using the ideas i do have.

That's just how humans usually work, and i don't see any way to improve it.

Though, if i'm right about this, then what's a proper application of AI in games?
That's easy to answer: Making games more dynamic, avoiding the restrictions of static stories and quest designed without the players contribution. For example advanced NPCs who generate their blah blah and animations dynamically and in response to the players actions.

I know one or two game developers who work on such things.
But they don't use AI tools to do so.
Those tools are for consumers, not for producers or creators. The latter ones make tools.
Imo it's important to understand the difference between those two things, for any creative industry.

Alberth
Alberth

I am not working in game-dev, but I do experiment with AI in programming context. I run the model locally, so it's far from state of the art due to processing limitations.

It works quite nicely for writing code for simple scripts. I had a problem with a number of xml files of an application, and I wanted to get some particular part of it, for all files. You can then say "Write a Python scripts that finds all XML files with extension .foo in a given directory. I want to extract information from it. For this example XML "...", it should produce this resulting Python data "...". " In 10-15 minutes I had the code to actually do it. Basically it saves you the time to have to manually type out the details of a straight-forward implementation.

Doing a bit of experimental scaffolding is nice too. It's quite easy to get a rough block of code for an idea/approach that you have. The general idea is usually there. I find this useful because with the code, I can then also see problems with the approach that I didn't see before. That enables me to make improvements in the idea or explore a different direction without having to write a full version first.

It can be useful in areas you are not familiar with at all. "I want to create a humanoid animation in 2.5D context" gave me several options to go about that problem, complete with suggestions of tools. I picked another high level solution, and it started explaining how to start with that in more detail, with example code to get started. I still have to try this, but it does give you nice pointers and starting points for more searching and reading documentation in very little time. Pretty likely it will not give me the results I like, but it looks like fun to try 🙂

At the other side, it sometimes doesn't do what I ask. "Translate this Lua code to Java" gave me "I deleted the non-relevant parts" as a highlighted point about the code that it wrote. One small problem though, there were no irrelevant parts 😛 (Perhaps it lacked context.)

It also tends to write lots of duplicated code, and spend cpu time recomputing values that you already have. This however very much depends on the input. For the duplicated code case, it also omitted a few functions which I then offered again. Then suddenly it suddenly understood to write a generic function instead of copy-pasting things. You can generally fix such things by specific requesting improvements.

In the details of the code, it makes semantic changes without telling me. I also found it may generate code that doesn't actually work because it violates an invariant of the data or an interface contract. So details need very careful checking.

Last but not least, I watched getting hopelessly lost in a coordinate-translating problem because it didn't truely understand the inter-play of values between "-x" and "x-y", combined with max/min functions. I also saw a smaller llm make clearly false claims about its code.

Alberth
Alberth

JoeJ wrote:

I know one or two game developers who work on such things. But they don't use AI tools to do so. Those tools are for consumers, not for producers or creators. The latter ones make tools. Imo it's important to understand the difference between those two things, for any creative industry.

I think that depends. I recently read in the news at this site that Innogames has been using AI for creating art in SunriseVillage (iirc) for a year, and nobody complained about it. It enabled them to reduce staff. Of course, they are in a market where development money is THE problem.

I have been playing that game a bit in the last year and I did notice not entirely correct positioning of some parts in the 3D graphics, but nothing really bad. It looked like their normal quality.

I eventually quit playing it because it randomly caused crashes of my computer system that needed a power cycle to fix. Probably it was something in the GPU drivers.

JoeJ
JoeJ

Alberth wrote:

It enabled them to reduce staff.

Well, maybe that's great. But that's the whole problem i see with AI.
AI replaces manual workers, so people have less money to keep the economy alive.
We could eventually tax the tech megacorps to support people having no jobs.
But the problem is: Those megacorps don't have any money either. Because to have to give away their KI tools for free, since obviously those tools have no value and nobody would pay for it.
The games created with AI have no value either. They are just generic slop, recycling the same old ideas which have been data mined from teh people who have now lost their job.

So who benefits in the end from AI?
So far it looks like absolutely nobody does. It just devalues everything. Literally and without any exception of statistical relevance.
We shoot our own foot still, even if we nowadays are seemingly too stupid to do it ourselves.
It's ridiculous.

Alberth wrote:

but nothing really bad. It looked like their normal quality.

Yeah, i see six finger hands less frequently.
Those image generators are way better than average artist skills, judging from things like realism, or anatomical correctness.
I have also heard some examples of AI music which are very good in terms of mood and composition. Some wrong notes, but it's pretty damn good regardless.

So why should anybody learn to draw, or to play an instrument, or programming video games?
Nobody will anymore. So the source of training data remains restricted to the works of people from the least few decades.
We become stuck, incompetent and trapped in a cycle of recycling our former greater self.

Nobody wins or takes advantage.
It's just disrespect about our own capabilities, loss of motivation, ambitions, and self confidence.

I can see how a majority lacking the skills feels empowered by AI.
But it's a trap. A fallacy.
It's much better to deliver something individual even if amateurish and bad, than some artificial perfect pattern.

But well. That's just my opinion. I'll resist and stick at it. Some need to, just in case...

RmbRT
RmbRT

JoeJ said:
We could eventually tax the tech megacorps to support people having no jobs.

It's not like they're paying taxes anyway, even now.

JoeJ said:
I can see how a majority lacking the skills feels empowered by AI.

Anyone claiming AI is doing a better job than himself should be embarassed and git gud. AI can't ever replace an expert. Just like those jogging AI robots look like they pooped their pants.

JoeJ said:
But the problem is: Those megacorps don't have any money either. Because to have to give away their KI tools for free, since obviously those tools have no value and nobody would pay for it.

The funny thing is, some people actually pay hundreds of dollars for AI. It's not like they actually get paid more for cranking out more code. So just for that weird perceived comfort, they give up a sizeable chunk of what would have gone towards monthly savings.

JoeJ said:
So why should anybody learn to draw, or to play an instrument, or programming video games? Nobody will anymore. So the source of training data remains restricted to the works of people from the least few decades. We become stuck, incompetent and trapped in a cycle of recycling our former greater self.

This is the final consequence of nihilism when also offered with the means allowing them to avoid any effort. Watched a documentation on London's old architecture recently. Literally even a sewer facility from a few hundred years ago looks like a royal palace. People strove for excellence and beauty and peak craftsmanship back then. Not for minimal viable products where functionality is the only concern, and beauty or soul is given no consideration.

BTW, I've been making some progress on my new graphics API specification, spent an entire day just prototyping its high level design. It's even less complex than OpenGL in its usage, because it removes the global state (which is also what makes OpenGL inherently single-threaded). And instead of constantly passing in values that have to be validated and translated into driver-native or HW-native data formats on every call, you just pass a value once and then you receive a handle to the prevalidated and transformed native value. And then for the actual commands you just pass the handle. This means the driver just has to check whether the handle is valid, and then it is already done with the validation, and already has all the transformed data. I don't like how they do it in Vulkan where the entire GPU pipeline state is one monolithic object that is immutable, and for each variation, you need to create a complete new object, which has to get fully verified and then gets cached on disk. Modern games often have gigabytes of pipeline caches, and services like steam have terabytes of shader & pipeline state caches in their cloud even for a single game, because they have to cache that anew for each platform & driver & HW combination.

So in my design, you basically have the same cached prevalidation, but only of individual values (like vertex attribute descriptors, texture parameter sets, blend config, rasteriser config, etc.). So you don't need those massively expensive shader pipeline permutation caches that you have to set up ahead of time or risk a stutter when you do them on demand. And you basically eliminate all the input validation from OpenGL.

Currently, I'm only working on covering the basic featureset of vertex & fragment shader based rendering. Later on, I'll extend this for compute, etc.. And also the API makes you just say whether you need to preserve the output of the frame as the input of the next frame, or whether you want to overwrite the entire canvas on the next frame. This removes the mandatory explicit clear, which on desktop costs something, but on tile-based GPUs, clearing is actually faster than not clearing, because it turns into a nop and also suppresses the initial fetch of the framebuffer contents from RAM. And then it also doesn't have to write back the depth buffer and other attachments to RAM, meaning you get a massive reduction in memory bandwidth on tile-based GPUs without VRAM.

Walk with God.
Tom Sloper
Tom Sloper

@st3f0n why did you not tell us that you're selling stuff, and how much you charge for your [experimental?] “Game dev ai prompt pack”? I mean, you start the post saying you've been experimenting, but then when someone follows your link they find you're selling your prompts. Mixed messaging, anyone?

-- Tom Sloper    --      sloperama.com
JoeJ
JoeJ

I don't like how they do it in Vulkan where the entire GPU pipeline state is one monolithic object that is immutable, and for each variation, you need to create a complete new object, which has to get fully verified and then gets cached on disk. Modern games often have gigabytes of pipeline caches, and services like steam have terabytes of shader & pipeline state caches in their cloud even for a single game, because they have to cache that anew for each platform & driver & HW combination.

Well, at least you can change constants in a shader on API side to create permutations.
For example, i often have the tree level as a compile time constant, and thus i need to create one permutation for each level. I guess i can avoid that this way. My GI project took half an hour to compile on older NV cards, so i surely want to.

But ofc. you still need all the pipelines. Imo the concept is fine because it enables to precompute lots of driver work.
It only became a problem when PS5 has eliminated loading times, so games now have to do it while playing.

I have predicted Steam would add such precompiled shader service years ago, and now it became a reality.
But that's only one option to solve it. The others are: Use less materials, or simply make better games so waiting for the load is worth it. : )

RmbRT
RmbRT

JoeJ said:
But ofc. you still need all the pipelines. Imo the concept is fine because it enables to precompute lots of driver work.

Enabling that is good. But forcing it is bad. If I just want to quickly boot up the thing to debug it, I don't want to do a 30 minute compilation beforehand. And maybe there are some materials or something that are rarely used and do need a minor unpredictable variation from some huge possibility space, and I would want the ability to cheaply instantiate those on the fly.

Do you perchance know whether OpenGL will also trigger these pipeline caching operations, or rather: is this inherently something that GPUs can't get around, or is it something simply demanded by Vulkan because they wanted to frontload all effort? Can an OpenGL driver achieve less stuttering / pipeline caching time compared to Vulkan? For example Sebastian Aaltonen proposed a different model here, where you don't cache the entire pipeline, but rather just parts of it, such as the blend state, rasteriser state, etc., and then the full pipeline description is assembled from these parts. That would massively cut down the state permutations, since the state is broken into smaller bits. Ideally, an API lets you do arbitrary state configurations at runtime without having everything fully pre-baked. But I also see the value in being able to pre-bake as much as you want.

I think some emulator got around the shader permutation problem by having an Übershader which was basically controlled entirely by uniforms, and it models the full customisability of the emulated GPU hardware, and then it uses that while an optimised / baked instance of it is getting compiled in the background. Because those emulators had the problem that there was basically an unpredictable and almost infinite set of GPU configurations that a game might invoke at any time, and it was just impossible to pre-bake them all, even for a single game.

I am also thinking about letting you record entire command lists, saving them as a handle, and then you can just replay that command chain by referring to its handle, and it no longer has to perform any of the validation on it. That might require freezing all involved resources such as buffer sizes, etc., but it could also be used to make your own complex multiDraw sequence basically, and only paying for the validation once.

So basically: allow to bake as much as you want, but also as little as you want, at a finer granularity, so that you can assemble a full pipeline from pre-baked parts, so that assembling the pipeline itself becomes fairly cheap, as all the individual components are already validated, but you don't have to necessarily rebake the entire pipeline just because one variable in one component changed. And then only the component where something changed has to incur the validation cost, and not all the unchanged parts. I assume that this should cut down the actual amount of prevalidation you need by many orders of magnitude, and make on-the-fly modifications so cheap that they become viable in the few cases where they are necessary.

Walk with God.
JoeJ
JoeJ

RmbRT wrote:

BTW, I've been making some progress on my new graphics API specification

I can tell you about the most fundamental flaw of GLSL, so you can avoid it...

GLSL lacks pointers. This prevents from reusing allocated integer LDS memory to hold float values later, and vice versa.
There are workarounds, actually uintBitsToFloat() and floatBitsToUint(). It's cumbersome, but it works.
Though, guess what - the genius API designers have forgotten to add such workarounds for 16 bit floats too, so i'm seriously fucked now.
Their omission also affects the option to store floating point numbers in shader storage buffers declared as integer data, so i'm fucked again, but at least i can fix this one with a lot of work on splitting my buffers into multiple.

In OpenCL, which uses C instead shitty shading languages, you can just cast your memory to whatever types as desired.

What were they thinking? Gfx programmers too dumb to use C, not knowing how bits of numbers map to memory, requiring foolproof shading languages such as glsl / hlsl?
They always make such silly mistakes for absolutely no reason. Thanks to that, i'll not achieve the ideal performance the HW (any HW!) could do otherwise. Yes i'm mad. I always knew some day this design flaw would hit me hard, and now it does.

However, afaik OpenCL can be compiled to Spir-V too, so Vulkan itself can do it properly. The issue is just the glsl language.
So i could eventually do some glslang hacking to add my own glsl keywords, or make my own shading language all together.
But i have never done such thing and lack experience, and because i want to sell my GI stuff, it's better to stick at the established shitty standards. >:(




RmbRT
RmbRT

JoeJ said:
I can tell you about the most fundamental flaw of GLSL, so you can avoid it... GLSL lacks pointers. This prevents from reusing allocated integer LDS memory to hold float values later, and vice versa.

yeah that's also what Aaltonen criticised heavily. Though the problem is, my graphics API initially has to be emulatable. I'm splitting it up into multiple feature set tiers, so that there is one tier emulatable by GL ES 2 / WebGL1, one tier that is emulatable by GL ES 3 / WebGL 2, and then I'll probably add a high tier that can do all that other stuff like pointers, etc.. The problem with raw access to the HW like with pointers is that you suddenly have to insist on specifics of the hardware's design at a very low level. So I would not want to force that into the base feature set of my API. The base feature set is basically just vertex+fragment shader rendering. Tier 2 is still vertex+fragment, but with support for integers in shaders, and for packed vertex formats like 2_10_10_10_REV, etc.. Tier 3 will then be vertex+fragment+compute. I don't think I'll add any HW raytracing support, though. For tier 3, I should probably add pointers. That also means I have to have my own shading language. Or maybe compute will be a separate feature you can enable, instead of having it in a strict onion hierarchy of feature tiers. Then I could add mesh shaders and stuff like that separately, without one implying support for the other, necessarily.

I think long-term, I'll have to fork the MESA project and directly implement an OpenGL variation in there, hopefully with minimal maintenance effort, so I'd only have to pull a new upstream version and it would ideally then work for new GPUs without any additional effort or merge conflicts or something.

My goal is to also be able to retrofit this graphics API onto old / simple hardware. Anything that can run OpenGL should also be able to benefit from my improved design, if a driver is ported. Otherwise, it will just fall back to API compatibility emulation using actual OpenGL underneath, with basically no overhead. And if a driver exists, it should reduce the CPU overhead of the graphics API by a lot. Not to Vulkan levels, but without making it as tedious to use as Vulkan.

I like how a high-level API like OpenGL does not prescribe HW details, which means you also don't have to care about it. However, that concept only works when you can convey enough intent through the high-level API so that the driver knows what you actually wanted to achieve, and can then leverage its knowledge of the HW characteristics to achieve exactly what you intended, in the most efficient way. And OpenGL failed to let you convey intent clearly, so it also just sucks. But the general level of abstraction that OpenGL resides at is just about right.

One example of such things is as I mentioned, how tile-based GPUs absolutely need a glClear() call at the start of the frame, otherwise, they are forced to read the old framebuffer contents from memory, modify them, and write them back. Same with the depth values and colour attachments. And the invalidate() calls for the attachments is also important there, so that those don't get written back to memory after the frame. But on a dGPU, the glClear() call has a real cost instead of saving you performance. The API should just let you tell it exactly what you want to achieve, not what specific steps it should take to achieve that. Because for a given hardware, the driver should know better than you what steps are necessary to achieve a desired outcome with the least amount of work.

JoeJ said:
So i could eventually do some glslang hacking to add my own glsl keywords, or make my own shading language all together. But i have never done such thing and lack experience, and because i want to sell my GI stuff, it's better to stick at the established shitty standards. >:(

Yeah. Either you make it emulatable / compatible with existing standards, or you have to roll your own modified driver. But since I plan to roll my own OS anyway, I can also do my own drivers. I just hope it's manageable somehow. And since my OS won't be compatible with existing software anyway, I don't have to worry about gradual adoption or anything. I've been saying this for years: computing needs a fundamental reset and be rebuilt from scratch, taking only the good parts of what we achieved, and breaking all backwards compatibility. It will take a few years to do that, but the final outcome will be much simpler, much more efficient, more stable and maintainable, etc..

JoeJ said:
What were they thinking? Gfx programmers too dumb to use C, not knowing how bits of numbers map to memory, requiring foolproof shading languages such as glsl / hlsl?

To be fair, I don't think OpenGL prescribes what kinds of floats the GPU actually has to have. The min specs for lowp, mediump and highp floats and ints are pretty bad:

Honestly, a union might be what you are looking for, from the sound of it. So that you can reuse memory, without caring about what exact bit pattern is in there. Though I assume basically all hardware that had compute shaders also had float32 and int32 as a supported data type.

Walk with God.
JoeJ
JoeJ

RmbRT wrote:

I don't think I'll add any HW raytracing support, though.

Yo also ignore all the other recent additions: Tessellation, mesh shaders, VRS, occlusion queries, and maybe some more.

Personally i share the opinion that pretty much all of this is specific crap and not generally useful, but...

RmbRT wrote:

My goal is to also be able to retrofit this graphics API onto old / simple hardware.

... maybe you are a bit too retro ; )

Making GPUs simpler is definitively the top priority. Imo, all those useless features from above have only one true purpose: Guarding the monopols of the (US) tech industry.

But to make this a successful and widely adopted move, newer GPUs must establish a new rendering method, so much better and more efficient nobody will miss those retro triangles and rays.

But you don't get there with a retro mindset. We need to research and find the new and better method. That's the primary objective, imo.
Now you may argue you want web games and all that, but - just recently i've played some FPS using gaussian splatting in the browser, and it worked well with great performance. GS is not yet what we want for games, But together with Dreams and Nanite, but also neural rendering, it just confirms the trend: We want something new, and the days of HW triangles are counted.

If i'm right about that, and the revolution is coming, you may actually waste your time on designing a better API for a retro standard(!)

RmbRT wrote:

I like how a high-level API like OpenGL does not prescribe HW details, which means you also don't have to care about it.

Notice that ALL of those details are just the result of iterating those rasterized triangles over decades.
So you're thinking about treating the symptoms, but not about the cure of the cause.

RmbRT wrote:

The API should just let you tell it exactly what you want to achieve

Absolutely impossible, because the API designers never know what i want to achieve and how.
So exposing just everything at the lowest level is the best thing they can do.
It's just hard to do so, if you have to cover multiple GPU architectures, from multiple vendors and platforms.

But this does not excuse the many totally unneeded mistakes in their designs.

RmbRT wrote:

But since I plan to roll my own OS anyway, I can also do my own drivers. I just hope it's manageable somehow.

Tbh, i can't imagine this is manageable to a single person, who has not even access to the proprietary specs of all GPUs out there.
Especially NV keeps everything secret. Maybe this has changed a bit in recent years, but i doubt it.

Choose wisely: Work on games or work on OS. There is not enough time for both, eventually.

RmbRT wrote:

I've been saying this for years: computing needs a fundamental reset and be rebuilt from scratch

Agreed, but this applies to almost anything else too, from game designs, up to governments.

The good news is: The rest is happening. Because everything we have is currently breaking down in real time.

The bad news is: If we mess it up again, our civilization will end here and most of us will die. This is our last chance. >:D Har, har!

RmbRT wrote:

Honestly, a union might be what you are looking for, from the sound of it. So that you can reuse memory, without caring about what exact bit pattern is in there. Though I assume basically all hardware that had compute shaders also had float32 and int32 as a supported data type.

No, please don't try to help with such ideas about abstractions / generalizations. That's exactly the cause of the problems, e.g. this obfuscated crap about 'highp, mediump, lowp'.

I will create shader permutaions as needed: One for f32, one for f16, support for floating point atomics gives another multiply. (I found out even my old Vega GPU already has the feature : )

I try to help myself, using a C preprocessor to avoid the need of writing many shaders. It works well enough. I don't need any help.

It's just that i can not achieve the best options due to those language issues. Finding the best possible compromise will be very difficult, timeconsuming, and in the end still disappointing. Just because some smart API designers tried to make it easy with restrictive and short sighted generalizations.





RmbRT
RmbRT

JoeJ said:
If i'm right about that, and the revolution is coming, you may actually waste your time on designing a better API for a retro standard(!)

Triangles need to stay, because that's how you do basic rendering, and it's more universally useful than the old sprite blitters that only supported 2D axis-aligned quads. If you wanted to render your OS's GUIs etc., you'd still use triangles. And it's simply more efficient to do that on the GPU than on a CPU. I'm open to other forms of geometric representation, though. I just haven't looked into them and so far, triangles are a fast path and simple enough for me to use.

JoeJ said:
Tbh, i can't imagine this is manageable to a single person, who has not even access to the proprietary specs of all GPUs out there.

Well, MESA is free software and it has acceptable driver support afaik. So I can just leech off their efforts, and would just need to repackage the functionality they already built, I guess.

JoeJ said:
Agreed, but this applies to almost anything else too, from game designs, up to governments. The good news is: The rest is happening. Because everything we have is currently breaking down in real time. The bad news is: If we mess it up again, our civilization will end here and most of us will die. This is our last chance. >:D Har, har!

Yeah, I think what's happening around us is actually an intentional collapse. But I think they already made their plans as to what they want to replace it with, and the fact that they need to make everything unbearable first, in order to be able to push their next new thing as a replacement, says something about how human-friendly that replacement will be.

JoeJ said:
Choose wisely: Work on games or work on OS. There is not enough time for both, eventually.

Yeah, it's just a side aspiration, I am obviously first going to make a game, and hopefully make enough money so that I can then focus on such leisures as finishing my compiler and building an OS and adapting the MESA drivers to my own API, etc..

JoeJ said:
But you don't get there with a retro mindset. We need to research and find the new and better method. That's the primary objective, imo.

I mostly care about ease of programming, and by that I mean, the ease of programming something that is highly efficient, not just some random slop. For now, even for the old way of rendering triangles, there isn't a good API. So I'll solve that first, and then we at least have a good starting point for that basic feature set. Lots of attempts to add new features to triangles were just flopped vaporware, but I think the basic featureset that has built up around triangles is just very solid and polished. Like having alpha-to-sample-coverage to get free blending and transparency (at least the 1-bit alpha kind) without the heavy bulk of blending, and it also works with MSAA, giving you anti-aliased, texture-defined outlines for shapes at no additional cost, and without making you lose your depth buffer. That kind of stuff is just great. And it's compatible with tile-based GPUs, too. Which is another reason why I want to keep triangles. For a low-power device, specialised hardware paths are more energy-efficient than manually programmed behaviour. So this keeps the minimum entry barrier for producing useful graphics hardware low.

JoeJ said:
No, please don't try to help with such ideas about abstractions / generalizations. That's exactly the cause of the problems, e.g. this obfuscated crap about 'highp, mediump, lowp'.

I think in cases where you legitimately don't care, it's alright. But when you do care, you should be allowed to insist on fp32 or fp16, for example.

JoeJ said:
Absolutely impossible, because the API designers never know what i want to achieve and how. So exposing just everything at the lowest level is the best thing they can do.

I agree you can never fully know what someone wants to achieve. But for example my CPU ISA idea: I want to have a pair of instructions that open a loop block, where the loop start instruction specifies how many loop iterations I want, and the loop end instruction simply returns to the start and increments. That way, the CPU knows exactly that the following code will run n times, and how often it already ran. It can keep the start of the loop pre-decoded in a cache, and make the loop end 0-cost. And, it can avoid the mandatory final branch miss when the loop ends, because you communicated to it exactly what you want to achieve. x86 can already do that with the REP MOVSB instruction, for example: it repeats N times and does not cause any branch misses or pipeline flushes when it ends. So the concept clearly works, but they simply didn't allow you to do the same for actual user-defined loops.

This kind of knowledge is not transferred in a low-level API which only gives you super-fine-grained commands. But I'm not saying I want to prevent the user from being able to give detailed instructions. I am just saying most of the workloads you have, let's say 99%, fall into a category where a single high-level command would allow the underlying hardware to do a much more efficient job than a tedious stream of dozens of fine-grained commands. For example to achieve the optimal data flow (eliminating framebuffer fetches via glClear() and writeback via glInalidateFrameBuffer()) on a tiled GPU, you'd have to explicitly query whether the GPU you're running on is tile-based, and then based on that, do those commands. And if you are on a dGPU, you want to avoid the unnecessary glClear() of the colour buffer, because you know you want to fill the entire screen with geometry anyway. So, suddenly, you as engine programmer have to care about some weird peculiarity of the hardware that the driver could just take care of for your, if you simply told it “I don't want to reuse last frame's framebuffer contents, and I also don't want to keep this frame's results for the next frame, either”. Then you don't have to do some fragile GPU model detection.

Update regarding the pre-caching: I made it so that you can specify which parts of the PSO (pipeline state object) should be baked, and which parts should be configured in a way that allows me to easily modify them. For example I might want to keep the vertex fetch description dynamic, allowing me to switch between many vertex layouts without switching to a new PSO. Of course, this generates more instructions in the vertex fetcher, but that's a trade-off that I can choose. Or I want to hardcode the vertex layout, but maybe want to dynamically switch between fetching a vertex attribute and using a constant fallback value, and it ends up as a boolean shader uniform and an if under the hood. Maybe that fallback value should get baked into the instruction stream, or maybe it should also be cheaply reconfigurable. So for basically every field of the PSO, you can set a bit flag of whether it should be permanently baked or reconfigurable. That lets you cut down on PSO permutations, at the cost of having some runtime logic or variables inserted into the bytecode. But that's now your choice to make. And then you can still bind a more specialised pipeline whenever you want. It's kind of how you had GL_STATIC_DRAW, GL_DYNAMIC_DRAW, and GL_STREAM_DRAW for vertex buffers, to inform the driver of the intended usage, so that it could do better memory management for you.

JoeJ said:
VRS

Assuming this means Variable Rate Shading: isn't that the same as doing MSAA and then stretching the framebuffer by 2x, effectively using the subpixel samples like real pixels?

Walk with God.
JoeJ
JoeJ

It seems Slang supports pointers to LDS memory. If i can cast them to other types this might save me.
...I may reduce my NV rant by a notable amount of 1% :D

RmbRT wrote:

Triangles need to stay, because that's how you do basic rendering

Triangles won't go away, but currently its questionable if its still worth to have specific HW rasterization units.
I mean, ofc we need to have that better alternative first. But i really want to get rid of that gfx pipeline complexity at any cost. : )

RmbRT wrote:

Yeah, I think what's happening around us is actually an intentional collapse. But I think they already made their plans as to what they want to replace it with, and the fact that they need to make everything unbearable first, in order to be able to push their next new thing as a replacement, says something about how human-friendly that replacement will be.

Oh, a conspiracy theorist?
Nah, they do not cause the collapse by intent. They are indeed just that stupid. :D

RmbRT wrote:

Yeah, it's just a side aspiration, I am obviously first going to make a game

Good so. But then, maybe you should work on the jump'n'run for fun, instead on gfx API abstractions?

RmbRT wrote:

I am just saying most of the workloads you have, let's say 99%, fall into a category where a single high-level command would allow the underlying hardware to do a much more efficient job than a tedious stream of dozens of fine-grained commands.

No. Because that high level command you talk about can only be about standard building blocks we use often.
Which implies we talk about building blocks from the past, since future building blocks are unknown at present.
So your stuff comes with an expiry date.

Imo, we don't need HW acceleration anymore for just that same reason. Its always a big win presently, but becomes bloat and burden pretty quickly. It hinders long term innovation more than it helps efficiency in the short run.

RmbRT wrote:

And if you are on a dGPU, you want to avoid the unnecessary glClear() of the colour buffer, because you know you want to fill the entire screen with geometry anyway.

Well, after you have implemented filling the screen with geometry, the nanoseconds spent on clearing the screen will feel so insignificant, you might as well decide to optimize that further, instead searching in Vulkan specs how to figure out if your GPU needs a clear() or not. :D

RmbRT wrote:

“I don't want to reuse last frame's framebuffer contents, and I also don't want to keep this frame's results for the next frame, either”

Iirc, you can tell VK just that. Idk if it also helps with 'doing only if actually needed on current GPU'. Likely not, since it tends to offload all responsibility to the programmer. So you need a data base of GPU configuration settings you obtain from testing.

And as said, you can also spend your time on optimizing real bottlenecks, for a noticeable benefit.

But you can not do every optimization, since time and money is finite.
So you need to optimize your own workflow too, not just the program. And this means saying 'i don't care about X'.
I mean, maybe it would have been better to work on low quality batteries, than on perfect fossil fuel engines?
German engineer precision is not always the guarantee to success ; )

RmbRT wrote:

For example I might want to keep the vertex fetch description dynamic, allowing me to switch between many vertex layouts without switching to a new PSO. Of course, this generates more instructions in the vertex fetcher, but that's a trade-off that I can choose. Or I want to hardcode the vertex layout, but...

So you are not 100% sure what you want for yourself alone, i see.
Personally, i just try to keep it simple by 'using just one vertex layout and material for almost everything'.
This way i have no problem with pipelines and all that.
Later i can add some more stuff for skinned meshes or hair reflections. Incremental additions, still manageable.

Maybe this is a problem if you make an engine for everybody. But as long as it's just for us, it won't ever become a big problem at all, no?

RmbRT wrote:

Assuming this means Variable Rate Shading: isn't that the same as doing MSAA and then stretching the framebuffer by 2x, effectively using the subpixel samples like real pixels?

Idk. I remember sebbie talking about those MSAA tricks a lot, but likely VRS goes further by allowing larger blocks at variable sizes. The user can provide an image to define how large the blocks are at which part of the screen.

Though, it's only a benefit for forward rendering i guess. For deferred you could easily implement this yourself in software.
(assuming VRS is only about shading, not rasterization as well)

RmbRT
RmbRT

JoeJ said:
I mean, maybe it would have been better to work on low quality batteries, than on perfect fossil fuel engines? German engineer precision is not always the guarantee to success ; )

Set one autist on each of those tasks. Both tasks solved!

JoeJ said:
Likely not, since it tends to offload all responsibility to the programmer. So you need a data base of GPU configuration settings you obtain from testing.

JoeJ said:
And as said, you can also spend your time on optimizing real bottlenecks, for a noticeable benefit.

So, the programmer using the API is constrained by time and wants to focus on hard tasks that he actually has to solve. Then why is the API also offloading the responsibility of identifying the GPU type & version to the programmer? When the driver for the API is written by exactly the people who manufactured the GPU? And the 1k lines until you can even draw a triangle… That's all time spent writing code that does not solve the actually hard problems of the project. That's why I mean I want to be able to just tell the driver who knows the HW best, just what I want to achieve, for things that I can expect the driver to be able to solve. And then for the game-specific problems, specific to the assets & environment and scenes and effects I want, not the HW-specific problem of how to most efficiently start & finish a frame, that's where I want to sink 99% of my time into.

JoeJ said:
So you are not 100% sure what you want for yourself alone, i see.

No, that was just a hypothetical. In OpenGL, that's standard behaviour. You have multiple VAOs, and they can all have their own vertex layout, and it's not a big deal. You can have vertex attributes that are only rarely actually used in the model, and most of the time are just a default constant value (like colour-tinting a model). So you may want vertex attributes that you can dynamically switch from a constant to being fetched from the vertex object.

JoeJ said:
Personally, i just try to keep it simple by 'using just one vertex layout and material for almost everything'. This way i have no problem with pipelines and all that.

True, de facto, I also only have one vertex layout per shader. But still, that doesn't mean the API itself should enforce that kind of design, and actively punish deviation from it, when earlier APIs like OpenGL could deal with it just fine. Deviating from some intended “idiomatic” way of using the API should not incur a cost solely because it's not what the API designers thought was best practice. It should only incur a cost if it's actually a technical necessity that's not a fault of the API designer. Having the ability to partially bake values, and to partially emit microcode for the GPU that can handle dynamic values at runtime, allows you to either opt into full vulkan-style baking, or into full OpenGL-style ad-hoc state management, or a mix of both, getting the performance of full baking for the parts of the pipeline that you requested to be static, and taking a slight performance overhead for the dynamic handling in other parts, but making the reconfiguration of those dynamic parts basically free. That way, you as the programmer only incur the costs that you want to incur, and retain all the freedoms you want to retain. That is what actual fine-grained control over the shader pipeline looks like. You, the programmer, are allowed to make informed choices about the exact behaviour of the pipeline even to the microcode level, without being bound by arbitrary “idiomatic” design choices by some purist API designer who hasn't ever written a game himself. You can still do the gigabytes of pre-baked PSO caches with my API. But you can also do completely ad-hoc PSO reconfiguration that is much cheaper than with Vulkan, but results in microcode that has to generically handle the pipeline state on the fly. And to quote your concern: when it actually becomes a bottleneck (the fact that the GPU microcode has to handle dynamic inputs), you can choose to fully bake a pipeline instead. By simply setting a few bit flags on the pipeline object.

JoeJ said:
Well, after you have implemented filling the screen with geometry, the nanoseconds spent on clearing the screen will feel so insignificant, you might as well decide to optimize that further, instead searching in Vulkan specs how to figure out if your GPU needs a clear() or not. :D

On an iGPU without VRAM, with laptop-tier memory bandwidth, and less L3 cache than the entire screen buffer takes, a clear call can easily invalidate your entire L3 cache and take a few microseconds. It's of course much cheaper on a dGPU with 128MiB of cache and much more memory bandwidth than an iGPU has. Evicting the entire L3 also means the CPU will take a performance hit, and the memory bandwidth used, also slows down all the cores, since the iGPU shares the memory bandwidth and cache with the CPU.

JoeJ said:
Maybe this is a problem if you make an engine for everybody. But as long as it's just for us, it won't ever become a big problem at all, no?

But for just me, even OpenGL 2.0 is sufficient, so that point is moot. I will never produce enough assets to meaningfully max out the performance of OpenGL 2.0, unless I do something ridiculous like infinite procedural worlds with infinite view distance. But for anything I can personally handcraft, it's unlikely I can even max out an iGPU. And if GL2 is not sufficient, then GL3 with VAOs could improve that again. Only if I did something dumb like having tons of dynamic lights in a scene, then I would need something like a compute shader to do F+ rendering with a light binning step. So that would require GL4 with compute. And even then, I would still not need Vulkan. And even if that doesn't suffice, I could start using indirect draw from the later GL4.x revisions.

This is not a practical discussion, but a matter of principles. The graphics API should have as little performance overhead as possible, even if the legacy API is already fast enough. For the same quality rendering code that I write, I want to get the maximum possible performance out of it, so the point at which I should start to optimise should only be dictated by how slow my code is, not by how much needless overhead the graphics API has. And so even something like frequent state modifications should be as cheap as it can be. Even if it's not the most effective way to use a GPU. But until I reach the theoretical limit of how fast I can get with that approach, I should not be forced to change the approach. Vulkan does not allow me to get anywhere near that limit, as far as I understand, while OpenGL does, as far as I understand. So if the GL API can make my naive rendering code have low-enough driver overhead so that pipeline modifications are not a bottleneck for my specific program, then it's fine if I write my code like that. Only when that becomes my actual bottleneck should I be forced to incur the cost and hassle (and development time spent) of baking pipelines. But when I need to, I should be able to, which OpenGL does not allow, but Vulkan does.

JoeJ said:
No. Because that high level command you talk about can only be about standard building blocks we use often. Which implies we talk about building blocks from the past, since future building blocks are unknown at present. So your stuff comes with an expiry date.

Imo, we don't need HW acceleration anymore for just that same reason. Its always a big win presently, but becomes bloat and burden pretty quickly. It hinders long term innovation more than it helps efficiency in the short run.

Again, battery-driven devices, embedded devices, phones, etc., they all want to be as power- & heat-efficient as they can be. Which means the more fixed-function HW you can use, the less power you consume, the more battery life you have, and the lower the risk of your phone catching on fire. It's like saying we don't need SIMD because scalar pipelining is so good. But yeah, if there is such a drastic paradigm shift that nobody uses triangles anymore and even how we think about managing framebuffers etc. becomes obsolete, of course it would no longer fit. In VR, for example, you need to render at 120+FPS on an embedded device that is also headmounted, so your heat emissions & power levels are super constrained. In such a case, you definitely want to avoid anything that would incur inefficiencies with the already limited HW capabilities you have. So just let the driver also solve some common problems that it can easily solve itself, but are tedious to properly solve on the programmer side. A small indie studio with 2 guys doesn't have the time to worry about making a comprehensive list of GPUs out there and their peculiarities, just so that the frame lifecycle can be optimised. When for the driver authors it's just a few lines of code. Vs. weeks spent by a small indie team to properly deal with all the vendors and devices and research them and all that, and then maintain and curate that list of devices over time even after release, as new devices come out, etc..

JoeJ said:
Triangles won't go away, but currently its questionable if its still worth to have specific HW rasterization units.

I'm fine with that. If a user's device is powerful enough to do rasterisation in software, why should I complain? A GPU is a blackbox to me anyway.

JoeJ said:
Good so. But then, maybe you should work on the jump'n'run for fun, instead on gfx API abstractions?

I'm writing that abstraction so that I can make my game & other programs target it, and then when used it in real projects, I can refine it, and finally some day implement a real driver for it (or rather, fork an existing driver and modify it to fit that API). This lets me identify pain points and strong points of the API design, in direct comparison with OpenGL, and in real-world scenarios. And since my games won't be graphically demanding enough to the point where the compatibility layer on top of OpenGL would make any noticeable difference, it's not harming, either. And it's just plain enjoyable.

JoeJ said:
I mean, ofc we need to have that better alternative first. But i really want to get rid of that gfx pipeline complexity at any cost. : )

But even if you have SW-driven rasterisation, you can't expect the API user to implement his own SW rasteriser. You still want to ship an SW rasteriser with the driver. Maybe with the option do also implement a custom one. But even then, you have the problem of whether you want to force shader permutations of the rasteriser shader, or whether you have a dynamically configurable rasteriser shader (via uniforms), and which parts should be configurable, and which ones should be baked via a hard permutation.

Walk with God.
JoeJ
JoeJ

RmbRT wrote:

Set one autist on each of those tasks. Both tasks solved!

If you look out for autists you came to the right place. I've met quite a few here already. One said they just like video games and computers, because it's easier to come along with that than with other people.
I recommend to never make any jokes about people with mental issues in general, but especially here in the games industry.
I mean, it's not their fault. Personally i do even think it would be good to get some guidelines on how to deal with affected people, to avoid all the dispute i have caused without intent. If i would run a large game studio, i would probably consider this. To me this seems really massive, but nobody talks about it.

However, so far none of them turned out to be a genius who can solve complex problems in no time. They can get lost in details easily, though.

RmbRT wrote:

So, the programmer using the API is constrained by time and wants to focus on hard tasks that he actually has to solve. Then why is the API also offloading the responsibility of identifying the GPU type & version to the programmer?

Because we asked for it. We were not happy about draw call overhead and complained, we were not happy about single threaded contexts and complained, etc.
And after decades of whining and begging, they finally listened, and they gave us Mantle, DX12, Vulkan.
It allowed us to solve our performance problems.
Ofc. this comes at the cost of doing more work. But we can not really complain about that. We asked for it.

Those not willing to do the extra work sticked at the older higher level APIs.
But as time went on, engine development was mostly outsourced to U-engines and experts, so low level just won and the old stuff was abandoned because nobody used it anymore.

Ofc. everybody tries to reduce the complexity by doing his own abstraction layers on top of gfx APIs, like you currently plan to do.
The good news is: You can do this much better on top of a low level API, since it gives you more options.

So it's actually fine as is, i would say.
But i still need to complain about issues i encounter, because there is always room for future improvement.

RmbRT wrote:

And then for the game-specific problems, specific to the assets & environment and scenes and effects I want, not the HW-specific problem of how to most efficiently start & finish a frame, that's where I want to sink 99% of my time into.

I hear you, but coming up with a good abstraction layer is MUCH more work than just writing those 10k lines of code once to draw a triangle.
And it may not matter much how precisely you start and finish a frame.
Just work on the game, and if performance problems show up, then you may need to dig in on certain subject details to optimize.
You don't need to solve any potential performance problem in advance. I mean, they call this 'premature optimization' and say it's a bid thing for a reason.

To be clear, i do not really think you tend to premature optimization. But rather, you may think that you can get things right from the start, with proper planning. But in my experience this rarely holds. I never get hard things right from the start. I come to some solution which solves the problem, but it usually turns out overcomplicated and inefficient. It may take years to simplify and optimize, and only after that time i know what the right way actually is, and ideally should have been from the start. That's just how it is all the time. So i begin with consciously doing it wrong, and i'm fine with that, because i just can't do better anyway.

Considering you want to model a jump with an analytical parabola, i'm afraid the same applies to you ; )

RmbRT wrote:

True, de facto, I also only have one vertex layout per shader. But still, that doesn't mean the API itself should enforce that kind of design, and actively punish deviation from it, when earlier APIs like OpenGL could deal with it just fine.

Sure, but it's exactly this freedom and ease of use which made OpenGL inefficient over time. Convenience overhead became a bottleneck.
I agree there is probably a better compromise in between low and high level extremes, but we have to work with what we get.
Just get used to the fact you have to build an entire pipeline for each variation. Build them at start up and call it a day. It does not cause you more work to do in practice.
I remember it was hard for me to accept i have to build two pipelines, one for each frame of a double buffered rendering setup. And if i use 3 frames i'd need to build 3 pipelines. This made - and still makes no - sense to me. But i had to do it.
(I can't remember if i've learned how to avoid this eventually. Maybe it's not really needed and there is a way around. I can't remember the solution if so, but i'll always remember the problem. Which illustrates it's maybe more a psychological than a technical problem ; )

RmbRT wrote:

But for just me, even OpenGL 2.0 is sufficient, so that point is moot. I will never produce enough assets to meaningfully max out the performance of OpenGL 2.0, unless I do something ridiculous like infinite procedural worlds with infinite view distance. But for anything I can personally handcraft, it's unlikely I can even max out an iGPU.

I notice you mostly focus on the problem to fill a screen with triangles. This was the primary problem in the 90's, when people focused on hidden surface removal, or the fastest way to raster triangles on CPU, etc.
But in present time the primary bottleneck is lighting. Low poly models are no guarantee you won't see any performance problems on iGPUs, i'm afraid.
Ofc. it's possible to make games without expensive or any lighting, but even with low ambitions it's usually the bottleneck.

RmbRT wrote:

so the point at which I should start to optimise should only be dictated by how slow my code is, not by how much needless overhead the graphics API has.

You can get there by using low level APIs. : ) With high level the overhead is just there and you can't do anything against it.

RmbRT wrote:

Again, battery-driven devices, embedded devices, phones, etc., they all want to be as power- & heat-efficient as they can be. Which means the more fixed-function HW you can use, the less power you consume, the more battery life you have, and the lower the risk of your phone catching on fire.

Yeah, mobiles are always the counter argument when i ask for flexibility over fixed function. And that's true.

But it's irrelevant. Phones do not even have buttons on them, so making a proper action game isn't possible already due to that.

I may need to change my mind only if in the future people can't afford PCs or consoles anymore, which is quite likely.
But till then those future mobiles might be powerful enough for proper games and some compute workloads i need for my non fixed function stuff. At least i hope so.

RmbRT wrote:

In VR, for example, you need to render at 120+FPS on an embedded device that is also headmounted

VR is a joke at this point. Even Zuck gave up on it.
The problem, beside obviously unsolvable motion sickness and vergence accommodation conflict, is:
The promise is true depth perception, achieved with image per eye.
But in practice, due to doubling the rendering cost, they fall back to non realistic lighting, so depth perception is way worse than looking on my realtime GI on a flat screen.

Imo, Metas mobile VR is bullshit, and Valves expensive room scale VR is bullshit too. Which is why it remains a small niche.
I would look into VR if there were some cheap glasses for PC, as initially intended by Palmer Luckey. This looked attractive to me, but seemingly was not good enough, sadly.

Real gaming still happens on PC and consoles, not on mobiles, VR goggles or web browsers.
All we need is a move from CPU + dGPU to APU, following how consoles do it. This way the PC platform can remain affordable.
I want a cheap little box for my gaming and computing needs, not a clumsy Switch handheld with built in screen and joypads just to add costs.
As a grown up i have no interest to play games while sitting in the train. And sadly, there are much more grown ups than little kids these days.

RmbRT wrote:

A GPU is a blackbox to me anyway.

But it isn't. It's a general purpose processor you can use for anything, even if less conveniently. And it's 50 times faster than your CPU for the same job.

RmbRT wrote:

But even if you have SW-driven rasterisation, you can't expect the API user to implement his own SW rasteriser. You still want to ship an SW rasteriser with the driver.

No. Render implementations are NOT the the job of GPU vendors. Just look at NVs DLSS 5 haluzinated faces, to see what happens if we take that bait. They should make chips, not more. The rest is our business. Imo it already goes too far that GPU vendors are now responsible for upscaling, which should be our job too.

If SW rendering becomes popular again (it already is because UE5 is used for >50% of games), then yes - there should be multiple implementations from multiple engine makers and game studios.
They should be individual and not all the same, so video game gfx become interesting and competitive again.
What does it help if all games look good but they also all look the same? Nothing. Because if there is no 'better', then there is no 'good' either.

RmbRT
RmbRT

JoeJ said:
However, so far none of them turned out to be a genius who can solve complex problems in no time. They can get lost in details easily, though.

I always assume when someone mentions autists in an oversimplified, but positive manner, that it's just a metaphor for someone who is obsessed with something he loves. Which, to be fair, is a commonly observable trait in autists. Just like German engineers aren't what they used to be anymore, either. German engineers also shouldn't feel discouraged by that positive stereotype being thrown around (or maybe they should, and improve their skills to match the expectations? I wouldn't mind if German engineering regained its glory).

JoeJ said:
No. Render implementations are NOT the the job of GPU vendors. Just look at NVs DLSS 5 haluzinated faces, to see what happens if we take that bait. They should make chips, not more. The rest is our business. Imo it already goes too far that GPU vendors are now responsible for upscaling, which should be our job too.

If getting started with even rendering a single triangle means you have to implement a vertex fetch shader, a pixel shader, a rasteriser shader that schedules pixel shaders and performs depth tests, alpha blending, stencil tests, etc., then you are suddenly at a lot of code that you don't want to write just to get started. And then you also need to implement the triangle clipper and all that stuff that nobody wants to deal with, because the exact expected behaviour is clearly established for decades already. And if you have to write your own rasteriser, that means you also have to do a really good job with the pixel shader job issuing and batching. Which means you have to know exactly what batch size is the best on any given machine it will run on. Which means you would have to update the game's code whenever a new generation of HW comes out, because you need to decide on the parameters for that HW. Having the ability to write it yourself is fine, but being forced to do it is bad. That's completely different from the AI upscaler & framegen crap going on. Those are non-essential / crutch features and only exist to hype up their AI bubble and to justify poorly optimised games, and boost sales because of vendor-locked tech.

JoeJ said:
But in present time the primary bottleneck is lighting. Low poly models are no guarantee you won't see any performance problems on iGPUs, i'm afraid. Ofc. it's possible to make games without expensive or any lighting, but even with low ambitions it's usually the bottleneck.

And shadows. Which is why games nowadays almost always use noisy shadows that need to get averaged out over time with temporal accumulation. Crimson Desert has horrific flickering artifacts in its lighting. Recently studied up a bit on stencil-based shadow volumes, they sound like a huge hassle for anything other than static lights + static shadow casters. Though the same geometry used for the shadow volume boundaries can also be used for godrays. Still only useful for static geometry, though.

JoeJ said:
But it isn't. It's a general purpose processor you can use for anything, even if less conveniently. And it's 50 times faster than your CPU for the same job.

I know. I just don't care enough to think about it in terms other than how I talk to it to render my humble triangle graphics. That's my relationship with GPUs.

JoeJ said:
If SW rendering becomes popular again (it already is because UE5 is used for >50% of games), then yes - there should be multiple implementations from multiple engine makers and game studios.

I like competition, but the entry barrier to participation should be feasible to achieve even for a single hobbyist who is just starting out. If I can't use the pure driver API to do graphics myself as a beginner, without being also an expert on software-driven rasterisation, then you end up in a future where only engine kiddies will exist ever again.

JoeJ said:
I hear you, but coming up with a good abstraction layer is MUCH more work than just writing those 10k lines of code once to draw a triangle.

But that's just the thing I love doing, no matter how long it takes.

JoeJ said:
You don't need to solve any potential performance problem in advance. I mean, they call this 'premature optimization' and say it's a bid thing for a reason.

I call it “non-pessimisation” instead. If it's no additional effort to do a better job from the start, why do the inferior thing? And if the API had that choice, it would be a no-brainer to do the non-pessimised thing just as easily as it would be to do the naive thing.

JoeJ said:
But i still need to complain about issues i encounter, because there is always room for future improvement.

Definitely. Well, I'll have a look at vulkan sometime, and I'll try to see whether I can get it to generate pipelines that can adjust their behaviour at runtime without requiring a hard permutation. But as far as I know, Vulkan intentionally does not allow that. If it did allow it, I'd have nothing against Vulkan, really, because you could just implement OpenGL in terms of Vulkan, then. But currently, if you do that, AFAIK, you get all these pipeline permutation hits at runtime that you would not get in a regular OpenGL driver that isn't using Vulkan under the hood.

JoeJ said:
But till then those future mobiles might be powerful enough for proper games and some compute workloads i need for my non fixed function stuff. At least i hope so.

LOL, recently saw a youtuber praise ARM→x86 emulation on phones, and how you could then play windows games on phones with a Proton-style compatibility layer. And then he said the GPU got to 80°C… Maybe good for cold winters when you're waiting at a bus stop and forgot your gloves.

———————————————————————————

Anyway, the discussion with you helped me think about my API design much more clearly, and I made a lot of progress over the course of this discussion. It still doesn't cover the modern scope of an API, but I think the design approach so far is flexible enough to allow you to either implement Vulkan with it, or to implement OpenGL with it (as the two extremes of choosing to either bake everything or bake nothing into the pipeline). I'd have to look at Vulkan (not going to read the 3000 pages spec though) to actually confirm that hunch, though. It's inherently multi-threading friendly, because the only state it has are the resource handles. But the commands themselves have no global state they modify. I don't really subscribe to your preference for graphics, but still, it was valuable input.

Walk with God.
JoeJ
JoeJ

RmbRT wrote:

(or maybe they should, and improve their skills to match the expectations? I wouldn't mind if German engineering regained its glory)

I count on you guys to build the worlds first fusion reactor just nearby me. \:D/

RmbRT wrote:

And then you also need to implement the triangle clipper and all that stuff that nobody wants to deal with, because the exact expected behaviour is clearly established for decades already.

I did all this beck then when working on software rendering, before GPUs took over.
It's pretty complicated, much more then the splatting i have currently in mind. It's also more complicated than ray tracing i would say.

But that's not the problem i see with triangle rasterization, which is: Triangles make LOD very difficult to solve, and textured triangles make it very difficult to cache lighting, so we end up recomputing it every frame for every pixel, which is just brute force and inefficient.
So even if rasterizing triangles is fast or good, what does this help if it prevents us from solving the other problems which have a much higher cost or optimization potential?

RmbRT wrote:

Which means you have to know exactly what batch size is the best on any given machine it will run on.

Choosing the ideal batch size is a problem you only have if can choose freely. But that's not the case for the pixels covered by a triangle. The triangle dictates how many pixels there are to shade.
But ofc. you will end up trying to process more triangles than just one, to keep most threads busy. That's probably one reason why we have those 2x2 pixel quads on GPUs which now cause issues with small triangles.
I don't see a perfect solution to rasterize triangles on a parallel processor with fixed thread group sizes. It's a compromise of spending more complexity for better saturation.
That's another reason why i like my splats. Each splat takes the same number of steps, no matter how large it is. I have very high thread saturation now, much higher than with the old ray tracing approach, which required some complex mechanisms to steal work from other threads and such stuff.
I just need to solve the 'undesired transparency' problem, but i have some ideas. At some point i want to continue this work, but not mow.


RmbRT wrote:

And shadows.

Well, shadows are ofc. already included in the term 'lighting'.

RmbRT wrote:

Which is why games nowadays almost always use noisy shadows that need to get averaged out over time with temporal accumulation. Crimson Desert has horrific flickering artifacts in its lighting.

I saw that. Really bad. But their displacement tech is interesting. It has heavy artifacts too, though.
But at least we see individual studios still innovating custom tech, which is what i expect to see.

Idk how they do shadows, but looking at UE, which does use stochastic shadows with TA, i see no related artifacts.
They have some stability issues in other areas, but their shadows seem very good.

RmbRT wrote:

Recently studied up a bit on stencil-based shadow volumes, they sound like a huge hassle for anything other than static lights + static shadow casters.

Dynamic stuff works just fine. You only need to find the edge loop between front and back-facing triangles seen from the light, and extrude those edges for the shadow volume. Which means you need edge adjacency in the mesh data structure, and ideally no insane geometry resolution. Finding such edges is usually done on CPU. There also were attempts to do it with geometry shader, effectively requiring to duplicate all triangles so the extrusion can be pulled out from any edge. But that's bad imo. I would still try to find edges from the mesh data structure with compute.

The real problem is probably the high fillrate required to draw the volumes.
For a game like Doom 3 the BSP helped to cull a lot of lights and geometry. And it was clear most triangles in a room can never cast a shadow. Such things don't hold for modern games, but i would like a new game reviving this old tech. It has its own distinct look hard to reproduce in other ways.

RmbRT wrote:

I like competition, but the entry barrier to participation should be feasible to achieve even for a single hobbyist who is just starting out.

Well, that hobbyist can just use game engines already now, and he'll never need to worry about gfx programming.

But yeah, for the hobbyist who wants to make a game from scratch, it became certainly harder. But he can still use older high level APIs, which i recommend for the first learning phase anyway. It's not a good investment in the future, but the things learned still apply with low level as well.

RmbRT wrote:

But that's just the thing I love doing, no matter how long it takes.

What? You like that, designing API abstractions? You enjoy it, eventually doing it all the time while brain is in idle mode?
Well, then i won't try to stop you. I was thinking i'd do you a favor by telling you that you do not have to torture yourself.

RmbRT wrote:

I'd have to look at Vulkan (not going to read the 3000 pages spec though) to actually confirm that hunch, though.

Sadly i can't tell you anything about Vulkan. I have already and again forgotten what a desicriptor is, i'm afraid ; )

RmbRT
RmbRT

JoeJ said:
I saw that. Really bad. But their displacement tech is interesting. It has heavy artifacts too, though.

I bet they could get rid of those artifacts around the screen edges if they just rendered some buffer zone outside of the screen as well. But yeah, screenspace geometry displacement is pretty neat, and much simpler than nanite, workflow-wise.

JoeJ said:
But that's not the problem i see with triangle rasterization, which is: Triangles make LOD very difficult to solve, and textured triangles make it very difficult to cache lighting, so we end up recomputing it every frame for every pixel, which is just brute force and inefficient.

what about per-vertex lighting (we have detailed geometry anyway these days), and saving the lighting in a transform feedback buffer? Or my idea to not just store per-vertex colours, but actually encode a 64-bit texture compression for each triangle, and then that is applied to the triangle. Then only the 64-bit value would have to be stored per vertex. Obviously, that still has a problem with large surfaces like walls. At least when the light is unmoved and the geometry is unmoved, then we have object-space light caching.

JoeJ said:
The real problem is probably the high fillrate required to draw the volumes. For a game like Doom 3 the BSP helped to cull a lot of lights and geometry. And it was clear most triangles in a room can never cast a shadow. Such things don't hold for modern games, but i would like a new game reviving this old tech. It has its own distinct look hard to reproduce in other ways.

I think it's definitely useful for static environments, with static lights. Either you pre-bake the cast shadow directly onto the surface geometrically, or you bake a shadow volume. Then when a dynamic object enters the light's radius, you can do shadow mapping, but only need to perform it on the dynamic geometry. All statically shadowed geometry will not need to do any shadow map lookups. Obviously also only worth it in some cases where the shadows being cast aren't super long. But that would give you pixel-perfect, sharp static shadows, and then resolution-dependent shadow resolution for shadows cast by dynamic objects or dynamic lights. And in cases where the shadow volumes would be too large, you simply switch to shadow maps entirely, and you can also bake the static shadow maybe. Or a mixture of all three approaches, where you sometimes bake the static shadow outline into the surface it gets cast on, via geometry, and when a shadow volume is small enough, you simply use that, and for anything involving dynamic objects or dynamic lights, you use shadow mapping. Then you can for each situation choose the tool that fits best.

JoeJ said:
Idk how they do shadows, but looking at UE, which does use stochastic shadows with TA, i see no related artifacts. They have some stability issues in other areas, but their shadows seem very good.

I heard it's because of their low ray counts that they reconstruct with AI.

JoeJ said:
But yeah, for the hobbyist who wants to make a game from scratch, it became certainly harder. But he can still use older high level APIs, which i recommend for the first learning phase anyway. It's not a good investment in the future, but the things learned still apply with low level as well.

That's why I want to make an API that is as efficient as possible, but also as 1-person-friendly as possible. And then I just make it flexible so that it can also do all the other stuff that the flagship APIs can do, but it's all opt-in, rather than mandatory. So you can start out with simple design and easy usage, and then slowly over time adapt the code to opt into more and more of the features that demand more complexity on the user side, like managing all the PSO permutations ahead of time, planning around the concept of baking, etc.. And then you can go as far as you want into Vulkan-like territory, one tiny step at a time, without switching APIs. And all it really takes to do that is to offer the ability to control what parts should be reconfigurable (by just setting a uniform or something similar) and what parts of the pipeline should be fully baked, requiring a full pipeline reconfig if you want to change them.

JoeJ said:
What? You like that, designing API abstractions? You enjoy it, eventually doing it all the time while brain is in idle mode? Well, then i won't try to stop you. I was thinking i'd do you a favor by telling you that you do not have to torture yourself.

I even enjoy designing a programming language and a CPU ISA and an OS / runtime API. But yeah I think about that stuff when I go to sleep or when I wake up. I also got a side project where I want to make an alternative web standard which doesn't need HTML, HTTP, JS, DNS, and all the complexity that comes with those. Like cross-site scripting vulnerabilities, cross-origin requests, “documents” being actually programs that spy on you and send your behaviour to some server, etc.. My Web spec would differentiate between documents and interactive applications, and even for something like interactive forms, it would sandbox the scripting so that it cannot track you. The part that handles page layout and styling should be completely trustworthy, so that you can run that without any worries of it being able to spy on you. I didn't get too far yet for the whole idea, but that's also in the pipeline for some later day.

Walk with God.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.