Seq2SeqScorer¶
- class minicons.scorer.Seq2SeqScorer(model: str | torch.nn.Module, device: str | None = 'cpu', tokenizer=None, **kwargs)¶
Bases:
LMScorerClass for Autoregressive or Incremental (or left-to-right) language models such as GPT2, etc.
- Parameters:
model – should be path to a model (.pt or .bin file) stored locally, or name of a pretrained model stored on the Huggingface Model Hub, or a model (torch.nn.Module) that have the same signature as a Huggingface model obtained from AutoModelForSeq2SeqLM. In the last case, a corresponding tokenizer must also be provided.
device (str, optional) – device type that the model should be loaded on, options: cpu or cuda:{0, 1, …}
tokenizer – if provided, use this tokenizer.
- add_special_tokens(text: str | List[str]) str | List[str]¶
Reformats input text to add special model-dependent tokens.
- Parameters:
text (Union[str, List[str]]) – single string or batch of strings to be modified.
- Returns:
Modified input, containing special tokens as per tokenizer specification
- Return type:
Union[str, List[str]]:
- encode(text: str | List[str]) transformers.BatchEncoding¶
Encode a batch of sentences using the model’s tokenizer. Equivalent of calling model.tokenizer(input)
- Parameters:
text (Union[str, List[str]]) – Input batch/sentence to be encoded.
manual_special (bool) – Specification of whether special tokens will be manually encoded.
return_tensors (str) – returned tensor format. Default ‘pt’
- Returns:
Encoded batch
- Return type:
BatchEncoding
- prepare_text(text: str | List[str] | transformers.BatchEncoding) Tuple¶
Prepares a batch of input text into a format fit to run LM scoring on.
- Parameters:
text – batch of sentences to be prepared for scoring.
- Returns:
Batch of formatted input that can be passed to
compute_stats
- conditional_score(prefix: str | ~typing.List[str], stimuli: str | ~typing.List[str], reduction: ~typing.Callable = <function Seq2SeqScorer.<lambda>>, prob: bool = False, base_two: bool = False)¶
Pooled estimates of sequence log probabilities (or some modification of it), given a prefix. Pooling is usually done using a function that is passed to the method.
- Parameters:
prefix (
Union[str, List[str]]) – a batch of prefixes or primes passed to the language model. This is what the sequence is conditioned on, and the model ignores the word probabilities of this part of the input in estimating the overall score.stimuli (
Union[str, List[str]]) – a batch of sequences (same length as prefix) that form the main input consisting of the sequence whose score you want to calculate.reduction (Callable) – Reduction function, is selected to be
lambda x: x.mean(0).item()by default, which stands for the avg. log-probability per token for each sequence in the batch.kw – model-specific keyword arguments to pass to the prepare_text function
- Returns:
List of floats specifying the desired score for the stimuli part of the input, e.g., P(stimuli | preamble).
- Return type:
List[float]
- prime_text_deprecated(preamble: str | List[str], stimuli: str | List[str], separator=' ') Tuple¶
Prepares a batch of input text into a format fit to run LM scoring on.
- Parameters:
preamble (Union[str, List[str]]) – Batch of prefixes/prime/preambles on which the LM is conditioned.
stimuli (Union[str, List[str]]) – Batch of continuations that are scored based on the conditioned text (provided in the
preamble). The positions of the elements match their counterparts in thepreamble.
- Returns:
Batch of formatted input that can be passed to
compute_stats
- distribution(batch: Iterable) torch.Tensor¶
Returns a distribution over the vocabulary of the model.
- Parameters:
batch (Iterable) – A batch of inputs fit to pass to a transformer LM.
- Returns:
Tensor consisting of log probabilies over vocab items.
- next_word_distribution(queries: List, surprisal: bool = False)¶
Returns the log probability distribution of the next word.
- compute_stats(batch: Iterable, source: Iterable, rank: bool = False, prob: bool = False, base_two: bool = False, return_tensors: bool = False) Tuple[List[float], List[float]] | List[float]¶
Primary computational method that processes a batch of prepared sentences and returns per-token scores for each sentence. By default, returns log-probabilities.
- Parameters:
batch (Iterable) – batched input as processed by
prepare_textorprime_text.rank (bool) – whether the model should also return ranks per word (based on the conditional log-probability of the word in context).
prob (bool) – whether the model should return probabilities instead of log-probabilities. Can only be True when base_two is False.
base_two (bool) – whether the base of the log should be 2 (usually preferred when reporting results in bits). Can only be True when prob is False.
return_tensors (bool) – whether the model should return scores as a list of tensors instead of a list of lists. This is important in some other convenient methods used in the package.
- Returns:
Either a tuple of lists, each containing probabilities and ranks per token in each sentence passed in the input.
- Return type:
Union[Tuple[List[float], List[int]], List[float]]
- sequence_score(batch, reduction=<function Seq2SeqScorer.<lambda>>, base_two=False, source_format='blank', source=None)¶
TODO: reduction should be a string, if it’s a function, specify what kind of function. –> how to ensure it is always that type?
- token_score(batch: str | List[str], surprisal: bool = False, prob: bool = False, base_two: bool = False, rank: bool = False, source_format: str = 'blank') List[Tuple[str, float]] | List[Tuple[str, float, int]]¶
- For every input sentence, returns a list of tuples in the following format:
(token, score),
where score represents the log-probability (by default) of the token given context. Can also return ranks along with scores.
- Parameters:
batch (Union[str, List[str]]) – a single sentence or a batch of sentences.
surprisal (bool) – If True, returns per-word surprisals instead of log-probabilities.
prob (bool) – If True, returns per-word probabilities instead of log-probabilities.
base_two (bool) – If True, uses log base 2 instead of natural-log (returns bits of values in case of surprisals)
rank (bool) – If True, also returns the rank of each word in context (based on the log-probability value)
- Returns:
A List containing a Tuple consisting of the word, its associated score, and optionally, its rank.
- Return type:
Union[List[Tuple[str, float]], List[Tuple[str, float, int]]]
- logprobs(batch: Iterable, rank=False, source_format: str = 'blank') float | List[float]¶
Deprecated since version Use:
compute_stats()instead.- Parameters:
batch (Iterable) – A batch of inputs fit to pass to a transformer LM.
rank (bool) – Specifies whether to also return ranks of words.
- Returns:
List of LM score metrics (probability and rank) and tokens.
- Return type:
Union[List[Tuple[torch.Tensor, str]], List[Tuple[torch.Tensor, str, int]]]