> For the complete documentation index, see [llms.txt](https://docs.goodfit.io/goodfit-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.goodfit.io/goodfit-docs/product/scoring-analysis.md).

# Scoring Analysis

Scoring Analysis allows you to automatically derive the optimal set of scoring rules to prioritise your dataset so that companies similar to a certain "Target" set have the highest scores and those that are not like them have the lowest.&#x20;

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2FzocnubTL9opT1D9uoAhh%2FScoring%20Analysis%20-%20Overview.gif?alt=media&amp;token=e1fa745b-5f1a-493d-ab21-812569d34338" alt=""><figcaption><p>Scoring Analysis</p></figcaption></figure>

To do this, we perform a 'correlation analysis' where we compare every GoodFit data point of all the target companies to other 'baseline' companies in your GoodFit dataset. Based on the frequency that certain data point values occur in each set of companies, we can work out which factors correlate more with the 'target' set than the baseline companies and use that to derive scoring rules automatically.

**Video run-through**

{% embed url="<https://www.loom.com/share/23bd155df1654110a0789dd8ca29399d?sid=badd083f-3ba9-4248-9626-fcf50a6630cd>" %}

### Getting started

To get started, navigate to scoring -> scoring analysis.&#x20;

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2F1CzrV8omE7Ds4t5lmqAK%2FScreenshot%202024-05-28%20at%2015.14.26.png?alt=media&amp;token=27e2f05d-405c-433b-ae28-0b36907bc4c8" alt=""><figcaption><p>Scoring Analysis page</p></figcaption></figure>

We allow you to create multiple scoring analyses, for example if you want to try different 'target companies' or have different scoring sets per team or region, or simply want to repeat the analysis over time and track the evolution over time.

You can create a new report by clicking "New scoring analysis". This presents you with a screen where you can name this analysis and paste a set of domains corresponding to the "Target" set of companies. These target companies would be the ones you are trying to make your scoring rank the highest. E.g. they may be won accounts or accounts that progressed far in the sales pipeline. The exact definition is flexible and may need to vary from business to business, however we recommend starting with your won accounts over a recent timeframe. Ideally we would have in the order of 50 companies, and we don't recommend using less than 20.

{% hint style="info" %}
These target companies need to be present in your GoodFit dataset as we perform the analysis based on your dataset. If they are not in your dataset, please contact us to add them for you.
{% endhint %}

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2FCPtLEbBJk8PQZkFIJoe9%2FSetting%20Up%20Scoring%20Analysis.gif?alt=media&amp;token=c0d8b8a7-e3ac-4098-8cf3-2f616450c1d1" alt=""><figcaption><p>Setting target account domains</p></figcaption></figure>

Note that we normalise the domains, so it's okay to paste web addresses directly. Once you click "Calculate scoring analysis" the system starts calculating the analysis. This can take a few minutes.&#x20;

### How to read a scoring analysis

Once the analysis is generated, it will display as a list of the mostly highly impactful data points (e.g. the data points that most influence the differentiation between the target companies and the baseline companies) and within each data point, a set of values and how that individual value correlates with the target companies and a suggested score for that value.

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2FpOtCUCQlGmC1VJWrKJ7L%2FScoring%20Analysis%20-%20Frequency%20Details.gif?alt=media&amp;token=616675b4-e837-4d24-869b-ec3192b492be" alt=""><figcaption><p>Scoring analysis report</p></figcaption></figure>

You can click into each data point and look at how the values correlate with the target companies vs the baseline companies, and therefore how impactful that value is to scoring. Hovering over the suggested score gives a breakdown of the frequency analysis of that value between the target companies and the baseline companies.

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2FiaPim2mPhidcwoyy3Xes%2FScreenshot%202024-05-28%20at%2015.34.13.png?alt=media&amp;token=9e0014fb-5ba9-42e5-8306-1034586131ec" alt=""><figcaption><p>Frequency analysis</p></figcaption></figure>

For example, in this case we can see this value ("factor") was present in 71% of the target companies, but in 8.6% across all others. Therefore we can reason as this factor is much more common in the 'target' companies its an important scoring rule, and we have therefore suggested a score of 10 for this value. We also provide a confidence score, which is an estimation of how confident we are that this is an impactful rule, and is decided based on how different the frequencies are and the number of times it occurs. See below for a more detailed explanation.

We automatically select the most impactful rules, however you can deselect some if you disagree, and select others if you think they should be included

{% hint style="info" %}
In order to create a score, you can select a maximum of 100 rules from this screen.
{% endhint %}

You can preview the scoring by clicking "Preview scoring" which shows the best and worst companies in your dataset based on those scoring rules. You should expect to see your target companies as the best fit along with similar companies.

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2F4gJFjpmSpfMaQiQyWhKg%2FScoring%20Analysis%20-%20Preview.gif?alt=media&amp;token=f5340f3d-3655-43cb-b95c-22ca01b7c7d8" alt=""><figcaption><p>Best and worst fit companies</p></figcaption></figure>

Once you are happy with the selected scoring rules, you can create a new score from it, which will add a new column to your GoodFit dataset with scores calculated using the suggested rules. Note that you can modify the rules and values as you see fit.

<figure><img src="https://1640582697-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FMPh5iOobj4telctyJWpa%2Fuploads%2FEIUZ3FVQFIIUHjMn5cov%2FScreenshot%202024-05-28%20at%2015.43.19.png?alt=media&amp;token=5e9bea28-5408-4997-a16d-01dcc33ba039" alt=""><figcaption><p>Saving a new scoring field</p></figcaption></figure>

### Scoring Analysis modes: "vs All" and "vs Other"

We offer two modes of Scoring Analysis:

* "vs All". In this case we perform the frequency analysis comparing the target "good" companies to all other companies in your GoodFit dataset.
* "vs Other". In this case we perform the frequency analysis comparing the target "good" companies to a supplied set of baseline "bad" companies.

In most cases, we recommend "vs All" as the purpose of scoring is to create a score that works to score any company in your dataset. However, sometimes you may wish to perform a more detailed analysis for a specific part of the pipeline. However care must be taken to interpret the findings in the context of those settings.

For example, it may be useful to analyse the set of "won" companies versus the "worked but lost" companies for a specific period. However this is only really useful to tell what's different about these two sets of companies, and this is unlikely to make a great score for your dataset as a whole.

### The details: How we perform the scoring analysis

We begin by dividing your GoodFit dataset into two 'sets': the 'good' set and the 'baseline' set. We then iterate over each data point that is of "picklist", "multi picklist", "boolean", "number", "percentage" or "currency" types. For each value within each data point (or range of values for numeric, see below) we perform a frequency analysis of that data point and value in the good and other sets.

For example, country=UK may occur in 40 out of 50 "good" companies, giving a frequency of 80%, whereas in the "baseline" set it may only occur in 300 out of 1000, giving 30%.

We use a "Two proportion Z-test" to work out the relative differences of these two frequencies based on the number of examples in each set and derive the confidence that the two frequencies are actually different given the number of occurrences. This gives us the confidence score and we use the confidence and difference between frequencies to derive the suggested score value.

Note that we treat data points differently depending on type:

* Picklist and Boolean. We consider each company as one example of this value
* Multi-picklist. We iterate the list of values per company and consider each occurrence as a count of this value.&#x20;
* Numeric fields. We apply a statistical bucketing algorithm to evenly spread the values out over a set of ranges, so we can work out the range that each companies value falls into and count these as discrete values.&#x20;

Note that we don't consider free text values as they tend to have unique values and have little impact on correlation.&#x20;
