Saturday, November 30, 2013

Measuring and validating the algorithm


In order to measure the performance and validate the algorithm described in the previous post, I created a validation script (attached below). The code randomly splits the character set in two and evaluates the similarity score. The assumption is that the two subsets of the same script should come out with a good similarity score (0 being identity - same character set). This is done 150 times for each of the scripts and outputs the average score received. The results are displayed below. For comparison, with this algorithm, the lowest similarity score between scripts is ~0.25 (between Thai and Gujarati) and the highest is ~25.24 (between Greek and Telugu).

Telugu: 2.17015555556
Cyrillic: 1.64583333333
Greek: 1.57917647059
Malayalam: 2.70133333333
Thai: 1.98205128205
Latin: 1.80256410256
Gujarati: 2.85166666667
Hebrew: 1.5647985348
Devanagari: 1.81041666667
Arabic: 2.29111111111
Tamil: 4.74801742919

The results are surprisingly good for such a simple algorithm (below 3 with the exception of Tamil). It seems to do better with the "straight - lined" scripts, such as Latin, Greek and Hebrew, than with the "round" scripts such as Tamil, Gujarati and Malayalam. I hope to improve the algorithm in the future, to reach a point where validation scores approach 0.


# -----------------------------------------------------------------------------
#
#  This script was created by Tamar Rucham
#
#  Measures the the algorithm in data_collection and languages_heatmap_results
#
# -----------------------------------------------------------------------------

from freetype import *
from data_collection import CalcChar, scripts

def run(chars, face):
    total_chars, total_contours, total_lines, total_curves = 0,0,0,0

    for singleChar in chars:
        ch = unichr(singleChar)
        contours, lines, curves = CalcChar(ch, face)

        total_chars = total_chars + 1
        total_contours = total_contours + contours
        total_lines = total_lines + lines
        total_curves = total_curves + curves

    return {'evarage_lines': float(total_lines)/float(total_chars), 'evarage_curves': float(total_curves) / float(total_chars)}

if __name__ == '__main__':
    import numpy
    import json
    import copy
    import random
    from languages_heatmap_results import getDiff


    face = Face('data/Arial Unicode.ttf')
    face.set_char_size( 48*64 )

    num_iterations = 150

    for scriptName, charsRange in scripts.items():
        total_diff = 0

        for iteration in range(num_iterations):
            diff = 0

            # Create two random subsets of the characters
            charsSlice1 = copy.deepcopy(charsRange)
            charsSlice2 = []
            for i in range(len(charsRange) / 2):
                # Dynamically get the len of the remaining of the array
                randIdx = random.randint(0, len(charsSlice1) - 1)
                charsSlice2.append(charsSlice1.pop(randIdx))

            slice1 = run(charsSlice1, face)
            slice2 = run(charsSlice2, face)
            total_diff += getDiff(slice1, slice2)

        print scriptName, ': ', (total_diff / float(num_iterations))

Wednesday, November 27, 2013

First algorithm to compare scripts

As I mentioned in my previous post from the FreeType python library I am able to get vector glyph information such as contours, lines and curves. For the first version of the algorithm for comparing scripts I created a python script to iterate over the character sets and output the average number of lines and curves per char. Then I wrote a second python script to grab that output and for each combination of scripts calculate the difference between them and output into a heatmap readable format. The formula is the total of the difference between average number of lines and average number of curves of the two scripts:

absolute_value(script_1_lines - script_2_lines) + absolute_value(script_1_lines - script_2_lines)

This produces a difference value of 0 between a script and itself and the higher the number - the bigger the difference. Both scripts are included at the bottom.

11 scripts that evolved from the Egypthian Hieroglyphs origin were analyzed in this manner and displayed in a heatmap with a tree relation to the left that demonstrates condensed evolutionary relationships between the scripts.



It is interesting to note that the script families do indeed form similarity "blocks". Though Arabic and Hebrew have more of a shared ancestry with the Brahmi family, it has a better geographical proximity to the Greek related scripts, it may also be related to the time of creation (of the scripts). The goal is to integrate such information in future versions.

The script to extract the data:

# -----------------------------------------------------------------------------
#
#  This script was created by Tamar Rucham, based on the FreeType library example 
#  glyph_vector.py
#
#  Assuming the ttf file and input file in a data folder from currently run script
#  this script will generate the statistics for each character in the given ranges
#  for the given languages, as well as general statistics for the language
#
# -----------------------------------------------------------------------------

from freetype import *
import numpy

# Extend the end of each script by 1 because of how ranges work
scripts = {
    'Latin': range(0x0041,0x005A+1) + range(0x0061, 0x007A+1),
    'Greek': range(0x0391,0x03A9+1) + range(0x03B1, 0x03C9+1),
    'Cryllic': range(0x0410,0x044F+1),
    'Hebrew': range(0x05D0,0x05EA+1),
    'Arabic': range(0x0621,0x063A+1) + range(0x0641, 0x064A+1),
    'Thai': range(0x0E01,0x0E2F+1) + range(0x0E40, 0x0E44+1),
    'Tamil': range(0x0B85,0x0B8A+1) + range(0x0B8E, 0x0B90+1) + range(0x0B92, 0x0B95+1)
            + range(0x0B99, 0x0B9A+1) + range(0x0BA3, 0x0BA4+1) + range(0x0BA8, 0x0BAA+1) + range(0x0BAE, 0x0BB9+1),
    'Malayalam': range(0x0D05,0x0D0C+1) + range(0x0D0E, 0x0D10+1) + range(0x0D12, 0x0D3A+1),
    'Telugu': range(0x0C05,0x0C0C+1) + range(0x0C0E, 0x0C10+1) + range(0x0C12, 0x0C28+1)
             + range(0x0C2A, 0x0C33+1) + range(0x0C35, 0x0C39+1),
    'Gujarati': range(0x0A85,0x0A8D+1) + range(0x0A8F, 0x0A91+1) + range(0x0A93, 0x0AA8+1) + range(0x0AAA, 0x0AB0+1)
             + range(0x0AB2, 0x0AB3+1) + range(0x0AB5, 0x0AB9+1),
    'Devanagari': range(0x0904,0x0939+1) + range(0x0958, 0x0961+1)
}

def CalcChar(singleChar, face):
    face.load_char(singleChar)
    slot = face.glyph

    outline = slot.outline
    points = numpy.array(outline.points, dtype=[('x',float), ('y',float)])
    x, y = points['x'], points['y']

    start, end = 0, 0
    lines, curves1, curves2 = 0, 0, 0

    # Iterate over each contour
    for i in range(len(outline.contours)):
        end    = outline.contours[i]
        points = outline.points[start:end+1] 
        points.append(points[0])
        tags   = outline.tags[start:end+1]
        tags.append(tags[0])

        segments = [ [points[0],], ]

        for j in range(1, len(points) ):
            segments[-1].append(points[j])
            if tags[j] & (1 << 0) and j < (len(points)-1):
                segments.append( [points[j],] )
        for segment in segments:
            if len(segment) == 2:
                lines+=1
            elif len(segment) == 3:
                curves1+=1
            else:
                # as reference - inner curves for complex curves
                for i in range(1,len(segment)-2):
                    A,B = segment[i], segment[i+1]
                    C = ((A[0]+B[0])/2.0, (A[1]+B[1])/2.0)
                curves2+=1
        start = end+1

    return len(outline.contours), lines, (curves1 + curves2)

if __name__ == '__main__':
    import json

    face = Face('data/Arial Unicode.ttf')
    face.set_char_size( 48*64 )

    languages_arr = []
    for scriptName, charsRange in scripts.items():
        print scriptName
        language_dic = {"language": scriptName, "chars": []}

        total_chars, total_contours, total_lines, total_curves = 0,0,0,0

        for i in charsRange:
            ch = unichr(i)
            contours, lines, curves = CalcChar(ch, face)

            total_chars = total_chars + 1
            total_contours = total_contours + contours
            total_lines = total_lines + lines
            total_curves = total_curves + curves
            char_dic = {"char": ch.encode('utf-8'), "contours": str(contours),
                        "lines": str(lines),"curves": str(curves)}
            language_dic["chars"].append(char_dic)

        total_chars = float(total_chars)
        language_dic["total_chars"] = total_chars
        language_dic["total_contours"] = total_contours
        language_dic["evarage_contours"] = (total_contours/total_chars)
        language_dic["total_lines"] = total_lines
        language_dic["evarage_lines"] = (total_lines/total_chars)
        language_dic["total_curves"] = total_curves
        language_dic["evarage_curves"] = (total_curves / total_chars)
        languages_arr.append(language_dic)

    outputFile = open('data/output.json','w')
    outputFile.write('{\n"languages":')
    outputFile.write(json.dumps(languages_arr,indent=4))
    outputFile.write('}')
    outputFile.close()

The script to calculate the heatmap data:

import json

def getDiff(char1, char2):
return (abs(float(char1['evarage_lines']) - float(char2['evarage_lines'])) + abs(float(char1['evarage_curves']) - float(char2['evarage_curves'])))

# def generateCharJson():
inputFile = open('data/output.json')
data = json.load(inputFile)
inputFile.close()

names_arr = []
data_arr = []
languages_data = []
maxData = 0
languages_info = data['languages']
row_index = 0
for language1 in languages_info:

lang1_name = language1['language']
names_arr.append(lang1_name)
row_arr = []
col_index = 0

# Now iterate through the rest of the chars and determine link weight
for language2 in languages_info:
lang2_name = language2['language']
weight = 0 if lang1_name == lang2_name else getDiff(language1, language2)
if weight > maxData: maxData = weight
row_arr.append([weight, row_index, col_index])
col_index += 1

data_arr.append(row_arr)

lang_data_dict = {lang1_name:[]}
for char in language1['chars']:
lang_data_dict[lang1_name].append(char['char'])

languages_data.append(lang_data_dict)
row_index += 1


outputFile = open('data/languages_heatmap.json', 'w')
outputFile.write('{\n"labels":')
outputFile.write(json.dumps(names_arr,indent=4))
outputFile.write(',\n"data":')
outputFile.write(json.dumps(data_arr,indent=4))
outputFile.write(',\n"languages_data":')
outputFile.write(json.dumps(languages_data,indent=4))
outputFile.write(',\n"minData":0')
outputFile.write(',\n"maxData":{}'.format(maxData))
outputFile.write('\n}')