You should simulate a number of repos and users to produce a realistic estimate. I would find that extremely useful, as would some of my colleagues at work. You should be able to scale up to a few hundred of each without too much trouble, and there will be no risk to running into any limits. I’m looking forward to this!
smallserverdata@lemmy.ml 3 days ago
This is the most useful comment in the thread and it is the thing I am going to build next.
You are describing the actual failure of what I posted. Hammering an HTTP endpoint with concurrent clients barely moves these apps. Forgejo went 171 to 313 MB under 24 concurrent clients at 4144 req/s, and the Go single binaries moved almost nothing, Caddy 40 to 47, ntfy 27 to 36. That is because the request path is cheap. What costs memory is data, so repo count and size and the working set of the database.
So the harness needs to create state, not traffic. What I plan for Forgejo is to create N repos through the API, push real history into them, create users, then measure at several values of N so you get a curve rather than one number. A curve is also more honest because your answer depends on your N.
Since you would use this: what values are worth reporting? I was thinking 10, 100 and 500 repos. And is repo count the thing that hurts, or is it total repo size, or CI, or concurrent git operations? You and your colleagues run this for real and I do not, so I would rather measure what you would actually check than guess.