---
title: Automated Machine Learning(AutoML) For Fast Data Labelling
description: This tutorial will help you understand that how automated machine learning (AutoML) is used to accelerate your labelling process.
image: https://blog.kili-technology.com/hubfs/Imported_Blog_Media/efficiency_comparison_with_without_model-2.png
---

![](https://cdn2.hubspot.net/hub/1598866/hubfs/logo-psdtohubspot.png?width=201&name=logo-psdtohubspot.png)

- Menu Item 1 
    - Sub-menu Item 1 
          - Another Item
    - Sub-menu Item 2
- Menu Item 2 
    - Yet Another Item
- Menu Item 3
- Menu Item 4

![](https://cdn2.hubspot.net/hub/459002/hubfs/Inbound_blog_.png?t=1471851895760&width=1440&name=Inbound_blog_.png)

# Hubspot Card Based Blog

Add Blog Subtitle's.

## Latest Stories

Search this site on Google

S

[Paul Graffan](https://blog.kili-technology.com/author/paul-graffan)

Jan 20, 2021

## [How To Monitor Machine Learning Model...](https://blog.kili-technology.com/blog/how-to-monitor-machine-learning-models-in-production)

[Non classé](https://blog.kili-technology.com/topic/non-classé)

<https://blog.kili-technology.com/blog/how-to-monitor-machine-learning-models-in-production>

[pierremarcenac](https://blog.kili-technology.com/author/pierremarcenac)

Dec 09, 2020

## [How Kili Technology and AutoML Helped...](https://blog.kili-technology.com/blog/automl-helped-scale-email-classification-for-customer-services)

[Non classé](https://blog.kili-technology.com/topic/non-classé)

<https://blog.kili-technology.com/blog/automl-helped-scale-email-classification-for-customer-services>

[Edouard](https://blog.kili-technology.com/author/edouard)

Dec 01, 2020

## [How to build a state of the art Machi...](https://blog.kili-technology.com/blog/how-to-build-a-state-of-the-art-machine-learning-platform-in-2021)

[Non classé](https://blog.kili-technology.com/topic/non-classé)

<https://blog.kili-technology.com/blog/how-to-build-a-state-of-the-art-machine-learning-platform-in-2021>

[Paul Graffan](https://blog.kili-technology.com/author/paul-graffan)

Sep 25, 2020

## [What is the best image segmentation t...](https://blog.kili-technology.com/blog/what-is-the-best-segmentation-tool)

[Non classé](https://blog.kili-technology.com/topic/non-classé)

<https://blog.kili-technology.com/blog/what-is-the-best-segmentation-tool>

## Featured Stories

Filter By Categories

- [Non classé](https://blog.kili-technology.com/topic/non-classé)
- [labeling](https://blog.kili-technology.com/topic/labeling)
- [kili](https://blog.kili-technology.com/topic/kili)
- [python](https://blog.kili-technology.com/topic/python)
- [API](https://blog.kili-technology.com/topic/api)

![](https://cdn2.hubspot.net/hubfs/459002/Author_Blog.png)

 By

[pierremarcenac](https://blog.kili-technology.com/author/pierremarcenac)

 May 07, 2020

<https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fblog.kili-technology.com%2Fblog%2Fautoml-fast-labeling> <http://www.linkedin.com/shareArticle?mini=true&url=https://blog.kili-technology.com/blog/automl-fast-labeling> <https://www.twitter.com/share?url=https%3A%2F%2Fblog.kili-technology.com%2Fblog%2Fautoml-fast-labeling> <https://plus.google.com/share?url=https%3A%2F%2Fblog.kili-technology.com%2Fblog%2Fautoml-fast-labeling>

## [AutoML for fast annotation](https://blog.kili-technology.com/blog/automl-fast-labeling)

  [Non classé](https://blog.kili-technology.com/topic/non-classé) [fast image annotation](https://blog.kili-technology.com/topic/fast-image-annotation) [fast text annotation](https://blog.kili-technology.com/topic/fast-text-annotation) [autoML](https://blog.kili-technology.com/topic/automl)

# AutoML for fast labeling with Kili Technology

This tutorial is taken from [our recipes](https://github.com/kili-technology/kili-playground). You can find an executable version of the [Jupyter notebook](https://github.com/kili-technology/kili-playground/blob/master/recipes/automl_text_classification.ipynb) on [Github](https://github.com/kili-technology/kili-playground).

In this tutorial, we will show how to use [automated machine learning](https://en.wikipedia.org/wiki/Automated_machine_learning) (AutoML) to accelerate labeling in Kili Technology. We will apply it in the context of text classification: given a tweet, I want to classify whether it is about a real disaster or not (as introduced in [Kaggle NLP starter kit](https://www.kaggle.com/c/nlp-getting-started)).

Why want to label more data when Kaggle often provides with a fully annotated training set and a testing set?

- Annotate the testing set in order to have more training data once you fine-tuned an algorithm (once you are sure you do not overfit). More data almost always means better scores in machine learning.
- As a data scientist, annotate data in order to get a feel of what data looks like and what ambiguities are.

But annotating data is a time-consuming task. So we would like to help you annotate faster by fully automating machine learning models thanks to AutoML. Here is what is looks like in Kili:

![](https://static.hsstatic.net/BlogImporterAssetsUI/ex/missing-image.png) ![](https://blog.kili-technology.com/hubfs/Imported_Blog_Media/automl-1.gif)

Additionally:

For an overview of Kili, visit [kili-technology.com](https://kili-technology.com). You can also check out [Kili documentation](https://cloud.kili-technology.com/docs).

The tutorial is divided into three parts:

1. AutoML
2. Integrate AutoML scikit-learn pipelines
3. Automating labeling in Kili Technology

## 1. AutoML

Automated machine learning (AutoML) is described as the process of automating both the choice and training of a machine learning algorithm by automatically optimizing its hyperparameters.

There already exist many AutoML framework:

- [H2O](http://docs.h2o.ai/h2o/latest-stable/h2o-docs/automl.html) provides with an AutoML solution with both Python and R bindings
- [autosklearn](https://automl.github.io/auto-sklearn/master/) can be used for SKLearn pipelines
- [TPOT](http://epistasislab.github.io/tpot) uses genetic algorithms to automatically tune your algorithms
- [fasttext](https://fasttext.cc) has [its own AutoML module](https://fasttext.cc/docs/en/autotune.html) to find the best hyperparameters

We will cover the use of `autosklearn` for automated text classification. `autosklearn` explores the hyperparameters grid as defined by SKLearn as a human would [do it manually](https://scikit-learn.org/stable/modules/grid_search.html). Jobs can be run in parallel in order to speed up the exploration process. `autosklearn` can use either [SMAC](http://ml.informatik.uni-freiburg.de/papers/11-LION5-SMAC.pdf) (Sequential Model-based Algorithm Configuration) or [random search](http://www.jmlr.org/papers/volume13/bergstra12a/bergstra12a.pdf) to select the next set of hyperparameters to test at each time.

Once AutoML automatically chose and trained a classifier, we can use this classifier to make predictions. Predictions can then be inserted into Kili Technology. When labeling, labelers first see predictions before labeling. For complex tasks, this can considerably speed up the labeling.

For instance, when annotating voice for [automatic speech recognition](https://en.wikipedia.org/wiki/Speech_recognition), if you use a model that pre-annotates by transcribing speeches, you more than double annotation productivity:

![](https://static.hsstatic.net/BlogImporterAssetsUI/ex/missing-image.png) ![](https://blog.kili-technology.com/hubfs/Imported_Blog_Media/efficiency_comparison_with_without_model-3.png)

## 2. Integrate AutoML scikit-learn pipelines

Specifically for text classification, the following pipeline retrieves labeled and unlabeled data from Kili, builds a classifier using AutoML and then enriches back Kili’s training set:

![](https://static.hsstatic.net/BlogImporterAssetsUI/ex/missing-image.png) ![](https://blog.kili-technology.com/hubfs/Imported_Blog_Media/automl_pipeline-1.png)

After retrieving data, [TFIDF](https://en.wikipedia.org/wiki/Tf%E2%80%93idf) pre-processes text data by filtering out common words (such as `the`, `a`, etc) in order to make most important features stand out. These pre-processed features will be fed to a classifier.

*Note:* `autosklearn` runs [better](https://automl.github.io/auto-sklearn/master/installation.html#windows-osx-compatibility) on Linux, so we recommand running code snippets inside a [Docker image](https://hub.docker.com/r/mfeurer/auto-sklearn/):

```
docker run --rm -it -p 10000:8888 -v /local/path/to/notebook/folder:/home/kili --entrypoint "/bin/bash" mfeurer/auto-sklearn
# cd /home/kili && jupyter notebook --ip=0.0.0.0 --port=8888 --allow-root
```

```
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.metrics import accuracy_score

MIN_DOC_FREQ = 2
NGRAM_RANGE = (1, 2)
TOP_K = 20000
TOKEN_MODE = 'word'

def ngram_vectorize(train_texts, train_labels, val_texts):
    tfidf_vectorizer_params = {
        'ngram_range': NGRAM_RANGE,
        'dtype': 'int32',
        'strip_accents': 'unicode',
        'decode_error': 'replace',
        'analyzer': TOKEN_MODE,
        'min_df': MIN_DOC_FREQ,
    }

    # Learn vocab from train texts and vectorize train and val sets
    tfidf_vectorizer = TfidfVectorizer(**tfidf_vectorizer_params)
    x_train = tfidf_vectorizer.fit_transform(train_texts)
    x_val = tfidf_vectorizer.transform(val_texts)

    # Select k best features, with feature importance measured by f_classif
    selector = SelectKBest(f_classif, k=min(TOP_K, x_train.shape[1]))
    selector.fit(x_train, train_labels)
    x_train = selector.transform(x_train).astype('float32')
    x_val = selector.transform(x_val).astype('float32')

    return x_train, x_val
```

Labeled data is split in train and test sets for validation. Then, `autosklearn` classifier is chosen and trained in a limited time.

```
from tempfile import TemporaryDirectory

# Un comment these lines if you are not running inside autosklearn container
# !conda install gxx_linux-64 gcc_linux-64 swig==3.0.12 --yes
# !pip install auto-sklearn

import autosklearn
import autosklearn.classification
from sklearn.model_selection import train_test_split

def automl_train_and_predict(X, y, X_to_predict):
    x, x_to_predict = ngram_vectorize(
        X, y, X_to_predict)
    x_train, x_test, y_train, y_test = train_test_split(
        x, y, test_size=0.2, random_state=42)

    # Auto-tuning by autosklearn
    cls = autosklearn.classification.AutoSklearnClassifier(time_left_for_this_task=200,
                                                           per_run_time_limit=20,
                                                           seed=10)
    cls.fit(x_train, y_train)
    assert x_train.shape[1] == x_to_predict.shape[1]

    # Performance metric
    predictions_test = cls.predict(x_test)
    print('Accuracy: {}'.format(accuracy_score(y_test, predictions_test)))

    # Generate predictions
    predictions = cls.predict(x_to_predict)
    return predictions
```

## 3. Automating labeling in Kili Technology

Let’s now feed Kili data to the AutoML pipeline. For that you will need to [create a new](https://cloud.kili-technology.com/label/projects/create-project) `Text classification` project. Assets are taken from Kaggle challenge `Real or Not? NLP with Disaster Tweets`. You can download them [here](https://www.kaggle.com/c/nlp-getting-started/data).

Connect to Kili Technology using `kili-playground` (Kili’s official [Python SDK](https://github.com/kili-technology/kili-playground) to interact with Kili API):

```
!pip install kili
from kili.authentication import KiliAuth
from kili.playground import Playground

email = 'YOUR EMAIL'
password = 'YOUR PASSWORD'
project_id = 'YOUR PROJECT ID'
api_endpoint = 'https://cloud.kili-technology.com/api/label/graphql'

kauth = KiliAuth(email=email, password=password, api_endpoint=api_endpoint)
playground = Playground(kauth)
```

Let’s insert all assets into Kili. You can download the original unannotated `test.csv` directly [on Kaggle](https://www.kaggle.com/c/nlp-getting-started/data).

```
import pandas as pd

df = pd.read_csv('./datasets/test.csv')
content_array = []
external_id_array = []
for index, row in df.iterrows():
    external_id_array.append(f'tweet_{index}')
    content_array.append(row['text'])

playground.append_many_to_dataset(project_id=project_id,
                                  content_array=content_array,
                                  external_id_array=external_id_array,
                                  is_honeypot_array=[False for _ in content_array],
                                  status_array=['TODO' for _ in content_array],
                                  json_metadata_array=[{} for asset in content_array])
```

Retrieve the categories of the first job that you defined in Kili interface. Learn [here](https://cloud.kili-technology.com/docs/projects/customize-interface/) what interfaces and jobs are in Kili.

```
project = playground.projects(project_id=project_id)[0]
assert 'jsonInterface' in project

json_interface = project['jsonInterface']
jobs = json_interface['jobs']
jobs_list = list(jobs.keys())
assert len(jobs_list) == 1, 'More than one job was defined in the interface'

job_name = jobs_list[0]
job = jobs[job_name]
categories = list(job['content']['categories'].keys())
print(f'Categories are: {categories}')
```

We continuously fetch assets from Kili Technology and apply AutoML pipeline. You can launch the next cell and go to Kili in order to label. After labeling a few assets, you’ll see predictions automatically pop up in Kili!

Go [here](https://github.com/kili-technology/kili-playground/blob/master/recipes/import_predictions.ipynb) to learn in more details how to insert predictions into Kili.

```
import os
import time
import warnings

!pip install tqdm
from tqdm import tqdm

warnings.filterwarnings('ignore')

SECONDS_BETWEEN_TRAININGS = 60

def extract_train_for_auto_ml(job_name, assets, categories, train_test_threshold=0.8):
    X = []
    y = []
    X_to_predict = []
    ids_X_to_predict = []
    for asset in assets:
        x = asset['content']
        labels = [l for l in asset['labels'] if l['labelType'] in ['DEFAULT', 'REVIEWED']]

        # If no label, add it to X_to_predict
        if len(labels) == 0:
            X_to_predict.append(x)
            ids_X_to_predict.append(asset['externalId'])

        # Otherwise add it to training examples X, y
        for label in labels:
            jsonResponse = label['jsonResponse'][job_name]
            is_empty_label = 'categories' not in jsonResponse or len(
                jsonResponse['categories']) != 1 or 'name' not in jsonResponse['categories'][0]
            if is_empty_label:
                continue
            X.append(x)
            y.append(categories.index(
                jsonResponse['categories'][0]['name']))
    return X, y, X_to_predict, ids_X_to_predict

while True:
    print('Export assets and labels...')
    assets = playground.assets(project_id=project_id, first=100, skip=0) ## Remove that
    X, y, X_to_predict, ids_X_to_predict = extract_train_for_auto_ml(job_name, assets, categories)

    version = 0
    if len(X) > 5:
        print('AutoML is on its way...')
        predictions = automl_train_and_predict(X, y, X_to_predict)

        print('Inserting predictions to Kili...')
        external_id_array = []
        json_response_array = []
        for i, prediction in enumerate(tqdm(predictions)):
            json_response = {
                job_name: {
                    'categories': [{
                        'name': categories[prediction],
                        'confidence':100
                    }]
                }
            }
            external_id_array.append(ids_X_to_predict[i])
            json_response_array.append(json_response)

        # Good practice: version your model so you know the result of every model
        playground.create_predictions(project_id=project_id,
                                      external_id_array=external_id_array,
                                      model_name_array=[f'automl-{version}']*len(external_id_array),
                                      json_response_array=json_response_array)
        print('Done.\n')
    time.sleep(SECONDS_BETWEEN_TRAININGS)
    version += 1
```

## Summary

In this tutorial, we accomplished the following:

We introduced the concept of AutoML as well as several of the most-used frameworks for AutoML. We demonstrated how to leverage AutoML to automatically create predictions in Kili. If you enjoyed this tutorial, check out the other Recipes for other tutorials that you may find interesting, including demonstrations of how to use Kili.

You can also visit the Kili website or Kili documentation for more info!

```

```

![](https://static.hubspot.com/final/img/content/email-template-images/placeholder_200x200.png)

### Popular Stories

<https://blog.kili-technology.com/blog/chatbot-training-datasets>

[36 Best Machine Learning Datasets for Chatbot Training](https://blog.kili-technology.com/blog/chatbot-training-datasets) Jul 07, 2020

<https://blog.kili-technology.com/blog/create-dataset-for-machine-learning>

[How to Create a Dataset to Train Your Machine Learning Applications](https://blog.kili-technology.com/blog/create-dataset-for-machine-learning) Feb 14, 2020

<https://blog.kili-technology.com/blog/dicom-medical-images-annotation>

[How to read & label dicom medical images on Kili](https://blog.kili-technology.com/blog/dicom-medical-images-annotation) May 27, 2020

### Popular Tags

[Non classé](https://blog.kili-technology.com/topic/non-classé) [labeling](https://blog.kili-technology.com/topic/labeling) [kili](https://blog.kili-technology.com/topic/kili) [python](https://blog.kili-technology.com/topic/python) [API](https://blog.kili-technology.com/topic/api) [Video annotation tool](https://blog.kili-technology.com/topic/video-annotation-tool) [annotating video datasets](https://blog.kili-technology.com/topic/annotating-video-datasets) [autoML](https://blog.kili-technology.com/topic/automl) [automatic video labelling](https://blog.kili-technology.com/topic/automatic-video-labelling) [callback](https://blog.kili-technology.com/topic/callback) [chatbot](https://blog.kili-technology.com/topic/chatbot) [computer vision applications](https://blog.kili-technology.com/topic/computer-vision-applications) [counterfactual](https://blog.kili-technology.com/topic/counterfactual) [data augmentation](https://blog.kili-technology.com/topic/data-augmentation) [data generation](https://blog.kili-technology.com/topic/data-generation) [dicom](https://blog.kili-technology.com/topic/dicom) [fast image annotation](https://blog.kili-technology.com/topic/fast-image-annotation) [fast text annotation](https://blog.kili-technology.com/topic/fast-text-annotation) [graphql](https://blog.kili-technology.com/topic/graphql) [how to create a dataset to train your machine lear](https://blog.kili-technology.com/topic/how-to-create-a-dataset-to-train-your-machine-lear) [how to create datasets](https://blog.kili-technology.com/topic/how-to-create-datasets) [how to train dataset](https://blog.kili-technology.com/topic/how-to-train-dataset) [image annotation categories](https://blog.kili-technology.com/topic/image-annotation-categories) [image annotation tools](https://blog.kili-technology.com/topic/image-annotation-tools) [image annotation types](https://blog.kili-technology.com/topic/image-annotation-types) [machine learning](https://blog.kili-technology.com/topic/machine-learning) [natural language processing](https://blog.kili-technology.com/topic/natural-language-processing) [nlp](https://blog.kili-technology.com/topic/nlp) [pneumonia](https://blog.kili-technology.com/topic/pneumonia) [pydicom](https://blog.kili-technology.com/topic/pydicom) [subscription](https://blog.kili-technology.com/topic/subscription) [text annotation](https://blog.kili-technology.com/topic/text-annotation) [text annotation example](https://blog.kili-technology.com/topic/text-annotation-example) [text annotation tools](https://blog.kili-technology.com/topic/text-annotation-tools) [text annotation worksheet](https://blog.kili-technology.com/topic/text-annotation-worksheet) [video annotation for deep learning](https://blog.kili-technology.com/topic/video-annotation-for-deep-learning) [webhooks](https://blog.kili-technology.com/topic/webhooks) [what is deep learning](https://blog.kili-technology.com/topic/what-is-deep-learning) [what is image annotation](https://blog.kili-technology.com/topic/what-is-image-annotation) [what is text annotation](https://blog.kili-technology.com/topic/what-is-text-annotation)

![](https://cdn2.hubspot.net/hub/2432204/hubfs/Quint-Assets/cta-banner.png?width=320&name=cta-banner.png "cta-banner.png")

![](https://cdn2.hubspot.net/hub/2432204/hubfs/Quint-Assets/cta-banner.png?width=320&name=cta-banner.png "cta-banner.png")

[Tweets by @inboundplace](https://twitter.com/inboundplace)

> Lorem Ipsum is a simple dummy text used as a dummy text contents. Lorem ipsum will be replaced. Lorem Ipsum is a simple dummy text used as a dummy text contents. Lorem ipsum will be replaced.Lorem Ipsum is a simple dummy text used as a dummy text contents. Lorem ipsum will be replaced.

[Previous Post](https://blog.kili-technology.com/blog/video-annotation-deep-learning)

##### [What is Video Annotation for Deep Learning](https://blog.kili-technology.com/blog/video-annotation-deep-learning)

[Next Post](https://blog.kili-technology.com/blog/webhooks)

##### [How to chain annotation projects with webhooks on Kili](https://blog.kili-technology.com/blog/webhooks)

[BUY *On* HUBSPOT](https://marketplace.hubspot.com/products/psdtohubspot/card-based-blog)

[![psdtohubspot](https://cdn2.hubspot.net/hub/1598866/hubfs/logo-psdtohubspot.png?t=1471851895760&width=201&name=logo-psdtohubspot.png "psdtohubspot")](http://www.psdtohubspot.com/)

©All Rights Reserved