Some of the examples are still wrong. Nuanced, but a Dutch native will still frown at it.
But more importantly is that you limited the context a lot. As in: the scope, the prompt, is very narrow.
In our case, we were generating emails. Lines like greetings are but one of 20+ details in that mail and not even the most important ones. The prompts ever larger, the multishot examples ever more tuned. And then, one in a few hundred will turn up with these "horrible" translations.
We've now moved to a chain of models, where we generate emails in American (the creative part) and then use another model to translate them to Dutch (the non-creative but culturally aware part). This works much better as we can pick models that are good at one thing or tuned to do this one thing better (either by the LLMAAS provider, or by parameters such as temperature).
With ChatGPT O1: https://chatgpt.com/share/67cfaa7e-70ac-8009-871b-571924b5a5... and with Claude's 'extended': https://claude.ai/share/4ba55410-98f3-4b53-9540-219acd2cdc4c