Gemini CLI
Run Google's Gemini CLI on ai.ml through its Gemini API base URL setting.
The Gemini CLI speaks Google's generateContent API. ai.ml serves it at /v1beta/models/{model}:generateContent, :streamGenerateContent and :countTokens.
With the CLI
npx @aiml/cli setup geminiBy hand
Export the two variables in your shell profile:
GOOGLE_GEMINI_BASE_URL="https://api.ai.ml"
GEMINI_API_KEY="aiml-live-…"They can also go in ~/.gemini/.env, but the CLI reads that file only in a folder it trusts. In any other folder it ignores the base URL, and would send your ai.ml key to Google.
Then set ~/.gemini/settings.json:
{
"security": {
"auth": {
"selectedType": "gemini-api-key"
},
"enableConseca": false
},
"model": {
"name": "gemini-2.5-flash"
},
"modelConfigs": {
"customOverrides": [
{
"match": {
"model": "base"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "flash"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "flash-lite"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "pro"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "auto"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "chat-compression-3-pro"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "chat-compression-3-flash"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "chat-compression-3.1-flash-lite"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "chat-compression-2.5-pro"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "chat-compression-2.5-flash-lite"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "chat-compression-default"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
{
"match": {
"model": "agent-history-provider-summarizer"
},
"modelConfig": {
"model": "gemini-2.5-flash"
}
}
],
"modelIdResolutions": {
"gemini-2.5-flash": {
"default": "gemini-2.5-flash",
"contexts": []
}
}
},
"agents": {
"overrides": {
"codebase_investigator": {
"modelConfig": {
"model": "gemini-2.5-flash"
}
},
"cli_help": {
"modelConfig": {
"model": "gemini-2.5-flash"
}
}
}
},
"experimental": {
"dynamicModelConfiguration": true,
"directWebFetch": true
}
}Each part of the settings file has a job:
modelConfigs.customOverridespins every model the CLI uses for its own helper calls to your model. Without it, those calls name other Gemini models: session titles at each start, chat compression, loop detection, sub-agents. ai.ml refuses any model not in its catalog.modelConfigs.modelIdResolutionswithexperimental.dynamicModelConfiguration. With an API key, the CLI renames any model whose name ends inflashto its newest flash model. This pair makes it keep the name you chose.experimental.directWebFetchfetches web pages without a model call. That model call would carry Google'surlContexttool, which only Google runs.security.enableConsecais off because the CLI's safety checker names its own model, and no setting redirects it.
Models
Name the model by its alias (gemini-2.5-flash), not its id. The Gemini API puts the model in one URL segment. Any catalog model works: the gateway translates, so claude-sonnet-4-6 is a valid model.name.
Verified with Gemini CLI 0.60.0.