How do you attribute GenAI serving cost across endpoints and experiments?
Several teams share model-serving capacity, and we need a fair way to show cost by application, endpoint, model version, and experiment. Request counts alone hide major differences in token volume and latency. Which usage tags, system tables, and allocation rules have helped you produce actionable chargeback data without adding too much instrumentation to every application?
0
