graphragzen.llm.openAI_API_client.OllamaClient

class graphragzen.llm.openAI_API_client.OllamaClient[source]

Interact with an LLM running on an Ollama Server.

!!During inference this class can force json output, but not the structure of the json.!!

The only reason this class exists is because the Ollama server does not have the ‘response_format’ feature; it can only be forced to output some json, but not force its structure. Forcing an output structure significantly increases the quality of generated graphs, concider using llama.cpp for serving your model.

Attributes

cache_persistent

chatnames

context_size

model_name

persistent_cache_file

tokenizer

use_cache

Methods

a_run_chat(chat[, max_tokens, ...])

Runs a chat through the LLM asynchonously

check_cache(llm_input)

Checks the hash(llm_in) -> llm_out cache and returns stored output if found.

format_chat(chat[, established_chat])

format chat with the correct names ready for tokenizer.apply_chat_template

num_chat_tokens(chat)

Return the length of the tokenized chat

print_streamed(stream[, timeit])

Streams the generated tokens to the terminal and returns the full generated text.

run_chat(chat[, max_tokens, ...])

Runs a chat through the LLM

tokenize(content)

Tokenize a string

untokenize(tokens)

Generate a string from a list of tokens

write_item_to_cache(llm_input, llm_output)

If a persistent cache file exists, this function can be used to append llm output to it.

__init__(*args, **kwargs)[source]

Interact with an LLM running on an Ollama Server.

!!During inference this class can force json output, but not the structure of the json.!!

The only reason this class exists is because the Ollama server does not have the ‘response_format’ feature; it can only be forced to output some json, but not force its structure. Forcing an output structure significantly increases the quality of generated graphs, concider using llama.cpp for serving your model.

Parameters:
  • base_url (str, optional) – url with API endpoints. Not needed if using openAI. Defaults to None.

  • model_name (str, optional) – Name of the model to use. Required when using openAI API. Defaults to “placeholder_model_name”.

  • context_size (int) – Context size of the model. Defaults to 8192.

  • api_key_env_variable (str) – Environment variable to read the openAI API key from. Defaults to “OPENAI_API_KEY”.

  • openai_organization_id (str, optional) – Organization ID to use when querying the openAI API. Defaults to None.

  • openai_project_id (str, optional) – Project ID to use when querying the openAI API. Defaults to None.

  • hf_tokenizer_URI (str, optional) – The URI to a tokenizer on HuggingFace. If not provided the API will be tested on the ability to tokenize. If that also fails a tiktoken is initiated.

  • max_retries (optional, int) – Number of times to retry on timeout. Defaults to 2.

  • use_cache (bool, optional) – Use a cache to find output for previously processed inputs in stead of re-generating output from the input. Default to True.

  • cache_persistent (bool, optional) – Append the cache to a file on disk so it can be re-used between runs. If False will use only in-memory cache. Default to True

  • persistent_cache_file (str, optional) – The file to store the persistent cache. Defaults to ‘./llm_persistent_cache.yaml’.

Return type:

None

async a_run_chat(chat, max_tokens=-1, output_structure=None, stream=False, **kwargs)[source]

Runs a chat through the LLM asynchonously

Parameters:
  • chat (List[dict]) – in form [{“role”: …, “content”: …}, {“role”: …, “content”: …

  • max_tokens (int, optional) – Maximum number of tokens to generate. Defaults to -1 (infinite).

  • output_structure (ModelMetaclass, optional) – Output structure to force. Ollama can only force some json, not the structure of the json. Making this non-empty forces a json output. Defaults to None.

  • stream (bool, optional) – Placeholder for compatibility with sync version, not used.

  • kwargs (Any) – Any keyword arguments to add to the lmm call.

Returns:

Generated content

Return type:

str

cache_persistent = False
chatnames: ChatNames = ChatNames(system='system', user='user', model='model')
check_cache(llm_input)

Checks the hash(llm_in) -> llm_out cache and returns stored output if found.

Parameters:

llm_input (str) – To check in cache for existing cached output.

Returns:

Union[str, None]

Return type:

str | None

context_size = 0
format_chat(chat, established_chat=[])

format chat with the correct names ready for tokenizer.apply_chat_template

Parameters:
  • chat (List[tuple]) – [(role, content), (role, content)] - role (str): either “system”, “user” or “model” - content (str)

  • established_chat (List[dict], optional) – Already formatted chat to append to. Defaults to [].

Returns:

[{“role”: …, “content”: …}, {“role”: …, “content”: …}]

Return type:

List[dict]

model_name: Any = None
num_chat_tokens(chat)

Return the length of the tokenized chat

Parameters:

chat (List[dict]) – in form [{“role”: …, “content”: …}, {“role”: …, “content”: …

Returns:

number of tokens

Return type:

int

persistent_cache_file = ''
print_streamed(stream, timeit=False)

Streams the generated tokens to the terminal and returns the full generated text.

Parameters:
  • stream (Iterator)

  • timeit (bool, optional) – If True display the number of tokens generated / sec. Defaults to False.

Returns:

Generated text

Return type:

str

run_chat(chat, max_tokens=-1, output_structure=None, stream=False, **kwargs)[source]

Runs a chat through the LLM

Parameters:
  • chat (List[dict]) – in form [{“role”: …, “content”: …}, {“role”: …, “content”: …

  • max_tokens (int, optional) – Maximum number of tokens to generate. Defaults to -1 (infinite).

  • output_structure (ModelMetaclass, optional) – Output structure to force. Ollama can only force some json, not the structure of the json. Making this non-empty forces a json output. Defaults to None.

  • stream (bool, optional) – If True, streams the results to console. Defaults to False.

  • kwargs (Any) – Any keyword arguments to add to the lmm call.

Returns:

Generated content

Return type:

str

tokenize(content)

Tokenize a string

Parameters:

content (str) – String to tokenize

Returns:

Tokenized string

Return type:

Union[List[str], List[int]]

tokenizer: Any = None
untokenize(tokens)

Generate a string from a list of tokens

Parameters:

tokens (Union[List[str], List[int]]) – Tokenized string

Returns:

Untokenized string

Return type:

str

use_cache = True
write_item_to_cache(llm_input, llm_output)

If a persistent cache file exists, this function can be used to append llm output to it.

Parameters:
  • llm_input (str)

  • llm_output (str)

Return type:

None