Saturday, December 21, 2013

Program – New Program Class

Currently a program is loaded by the main window class and given to the edit box using the set plain text access function.  This will set the text document of the edit box, which causes a document changed signal.  When this signal is processed, the edit box updates the program unit that is attached.  Because the text cursor is not valid during this operation, the recreated text from the program model cannot be inserted into the text document (the first of two issues mentioned in the last post).

One way to resolve this issue is to load the program directly into the program unit bypassing the edit box.  The program generates signals that lines have been changed by line number.  The edit box receives these signals, retrieves the recreated text for the lines from the program and puts them into the document.  Special handling is needed for lines that contain errors since there would be no code for these lines to recreate text from.

Eventually, when support for subroutines and functions is added, a loaded program will consist of a number of program units, one for the main routine and one for each of the subroutines and functions.  The program will consist of a list of program units.  When a program is loaded, the load routine will create a new program unit for each routine or function.  Likewise for the save routine, which accesses the list of program units.

The list of program units will be contained in a new program class.  This class will contain the routines for loading and saving programs.  The current program file path name will also be contained in this class.  New source and header files were added for the new program class.  To start simple, the new class only contains the program file path.  Since the file path is included in the applications settings, save and restore settings routines were also implemented.

[commit 17f3e958ed]

Sunday, December 15, 2013

Edit Box – Recreated Line Replacement

When a program line is changed or inserted, the line should be recreated and the recreated text should replace the line entered in the edit box.  To accomplish this, the program model class will send a signal with the line number of a line that is changed or inserted.  The edit box will receive this signal, retrieve the recreated text for the line, and replace the text of the line with this recreated text.

The new program changed signal was added to the program model class.  The update line routine was modified to send this signal when a line is changed or inserted.  When a line is changed, the actual code of the line may not have changed, for example, if spaces were added or removed, or the case of keywords was changed.  In this case, the program code will not be modified, but this signal still needs to be sent so that the edit box reflects the correct recreated program code.

The new program changed slot routine was added to the edit box class, which starts be retrieving the recreated text for the changed line.  A text cursor is obtained for the edit box and its position is set to the beginning position of the block containing the line.  The cursor is moved to the end of the block keeping the anchor at the beginning, which selects the entire line.  The recreated text is inserted at the cursor and since text is selected, the selected text is replaced.

The program model line text routine was modified to return a null string if the line has an error.  The program changed slot routine does not replace the text if the text is null indicating that the line has an error.  This required the recreator class recreate routine to be modified where the output string is initialized to an empty string (a pair of double quotes) instead of being cleared.  Clearing a string creates a null string, and a null string is not quite the same as an empty string (a null string is empty but an empty string is not null).

When text is replaced in the document using a text cursor, document changed and cursor moved signals are generated from the document.  These signals need to be ignored when the line is being recreated, otherwise an infinite loop occurs because the document change signal updates the program, which generates another program changed signal, an so on.  A flag was added to the edit box and is set before replacing text and cleared afterward.  The document changed and cursor moved slot routines were modified to do nothing if this flag is set.

There are two unresolved issues resulting from these changes.  The first issue occurs during the initial loading of a program.  As a new program is being loaded, each line added to the program should be recreated to the edit box.  However, this cannot occur because until the program has been loaded into the document of the edit box, the text cursor is not valid, so can't be used to replace text.  The second issue occurs when a line is replaced with recreated text; extra undo commands are added to the undo stack.

[commit 4b34bd2dde]

Saturday, December 14, 2013

Edit Box – Lines Changed Signal

The edit box was using the lines changed signal to notify the program model instance when lines changed in the document that included the starting line number, the number of lines deleted and the number of lines inserted and the text of the lines.  The program model update slot routine was connected to and processed this signal.  Since the edit box can access the program model directly, this signal is not needed.  The tester class is already accessing this routine directly.

The lines changed signal was removed and the two emits of this signal were changed to calling the update routine directly.  The update routine was changed from a public slot to a normal public member function in the program model class.  A few comments were added and word 'slot' was added to the comments of all the slot functions so that these would be easier to identify.

[commit c864bdde07]

Program – Error List Handling

The program model class was keeping a list of any errors detected when lines are translated.  After a group of line changes were processed, if the list of errors changed, the entire error list was sent to the edit box via a signal.  The edit box stored this list of errors and used it to generate its extra selection list, which is used to highlight the errors.  Since the edit box now has access to the program model, it was no longer necessary for the edit box to keep a copy of this list.

So that the edit box can keep its extra selection list up to date, the program model was modified to send signals for when an error item has been inserted, changed, or removed.  The connected edit box slots will update the extra selection list accordingly.  When the program model is done updating the error list, it send a signal that the error list has changed.  The connected edit box slot sets the extra selections to the base QPlainTextEdit class from updated extra selection list.  The edit box no longer scans through the error list to generate its extra selection list.

The functionality for finding the next or previous error from the current cursor location was moved from the edit box class to the program model class.  The current line number and column are passed to these routines, which are set to the next error line number and column upon returning along with a flag of whether the end or beginning of the program was passed so that a message can be issued.

A new routine was added to the program model class to handle when the current line is edited and contains an error.  The error is shifted if the edit takes place before the error, or deleted if the edit takes place within the error.  If the error list changes, the appropriate signals are sent to the edit box.  This routine was not made a slot since the edit box can call it directly.

During the initial loading of the program, the text cursor in the edit box is not valid.  Any errors in the program are sent as inserted error signals.  Since the text cursor is not valid, the extra selections cannot be created (each extra selection contains a format and a cursor, which needs to be set to a valid cursor).  The errors are temporarily saved in a list.  After the text cursor  becomes valid, extra selections are created from this list (which is then cleared).

The error list class used by the program model to hold the list of errors previously kept track of the first and last index affected by changes to the list.  These indexes were used by the edit box to maintain its extra selection list.  With the new direct change signals, these indexes are no longer needed, so the error list was changed to having a simple changed flag.  The error list changed signal is only sent when this flag is set.  The has changed access function was modified to clear this flag after it is read so a separate reset function is not needed.

Finally, the main window class status bar update slot was modified to receive the error message by an argument in the signal instead of it retrieving the message for the current line from the edit box, which retrieved the message from its copy of the error list.  The routine sending this signal in the edit box class was modified to send the error message, which is retrieved from the program model via its list of errors.

[commit 19c5f06d0c]

Tuesday, December 10, 2013

Edit Box – Program Unit Access

The edit class box needs access to the program model, specifically to the program unit currently opened in the instance of an edit box.  Eventually when subroutines and functions are implemented, any one of them could will be opened with their own edit box.  There could be several edit box instances opened at any given time.  Right now there is just a single program unit, the main routine, which will be opened in a single edit box instance.

To allow an edit box instance access to its program unit, a new pointer to a program unit instance (program model class) was added to the constructor of the edit box class.  This pointer is stored in a new program unit pointer member variable.  The connection of the line changes signal (from the edit box class), and the error list changed signal (from the program unit) are now made in the constructor of the edit box instead of the constructor of the main window class.

The constructor of the main window class was modified to create the program model first (the program unit for the main routine) before creating the edit box instance, which now requires the pointer to program unit instance that it will be editing.

[commit aa2125609a]

Saturday, December 7, 2013

Program (Recreator) and GUI Integration

The recreator is fully integrated with the program model such that program lines can be converted back into text.  When lines are entered into the program, the lines need to be recreated back to text and put into the edit box (the GUI), specifically into the text document of the edit box.  Like the temporary program view (being used for debugging) is the viewer of the data held by the program model, the edit box is the viewer of the data contained in the document.  The edit box also allows editing, so it is more than just a viewer.

The document of the edit box is really just the text representation of the program.  The program model holds the actual data of the program.  Ideally, the program model would be the document of the edit box and it would convert text to program code and back while editing.  However, Qt does not have an abstract text document class from which a document sub-class could be built that would hold its data in another form like program code.  The QTextDocument class is meant for text.

Alternatively, a new viewer could be designed that would allow all the text editing features (cut, copy, paste, undo, redo, etc.) like the QPlainTextEdit class that the edit box class is based on.  Designing one would be quite an effort.  Therefore, the edit box will the viewer for two data models at the same time, the text document (to allow text editing) and the program model (for holding the program code).  The program model will be the master of the data, with the text document being updated as the program changes.

This implies that the edit box either own the program model with the program code or at least have easy access to it like via a pointer.  The later approach will be used since the main window class will ultimately be the owner of the program.  Eventually there will be a list of program models, one for the main routine and several for the subroutines and functions of the program.  There will only be associated edit box instances when the main routine, subroutines or functions are open for editing.

Since the edit box will now have access to the program model, signals from the program model (for program changes) do not need to contain actual data.  For instance, when a program line has changed, its recreated text is needed to update the text document.  The signal could contain both the line number and text (already recreated).  Looking at the edit box to document interface, when the document changes, only the position, number of characters removed and inserted are contained in the signal.  The edit box must obtain the actual text changes by querying the document.  So, when the program changes, only the line number will be sent and the edit box will request the recreated text from the program model.

Wednesday, December 4, 2013

Information Dictionaries – Improved Design

The design of the info dictionaries required the program model to create the instances for the additional info for the dictionary (in this case, the constant number and string dictionaries) and pass this instance for the creation of the dictionary.  The program model owned and was responsible for these instances.  This is possibly problematic because it did not guarantee that an additional info instance of the correct type was passed to the info dictionary.

The design was changed where new constant number and string classes derived from the base info dictionary class were added.  In their constructors, the additional info instance is created of the correct type and they own the instance in the abstract info member pointer in the base class.  A destructor was added to the base class to delete this instance, which required a virtual destructor in the abstract info class so that the derived info class destructor gets called.

While not needed until the run-time module is implemented, access functions for the arrays in the additional info of the constant string and number dictionaries were added.  These functions simply call the access functions in the derived info classes.  However, a type cast to the derived class is needed since the base class defines the additional info instance pointer as an abstract info class pointer.

[commit 451edd6346]

Saturday, November 30, 2013

New Information Dictionary – Implementation

The information dictionary class was changed to a normal class derived from the dictionary class. The constructor is given a pointer to the information class instance created outside of the dictionary.  This pointer is saved in a pointer defined as an abstract information pointer, which can hold any information class pointer derived from the abstract class.

The add routine first adds the dictionary entry by calling the base dictionary class add routine with the token and case sensitivity option along with a pointer to the new entry flag so that it knows if a new entry was added, a removed entry was reused or an entry already exists.  If a new entry was added, an element is added to the additional information by calling the add element interface function of the information instance.  If the entry did not exist, the addition information is set from the token by calling the set element interface function of the information instance.  The index is returned.

The remove routine first removes the reference to the dictionary entry by calling the base dictionary class remove routine for the index specified.  If the entry was removed because it is no longer used, then the additional information for the element is cleared by calling the clear element interface function of the information instance.  The base dictionary class remove routine was modified to return whether  the entry was removed (made available to reused) or not.

The abstract information class defines the interface to the additional information.  The functions are defined as virtual functions with no default functionality so that derived information classes do not implement a function that it does not need.  The constant number and string information classes were changed from holding just a single element to being derived from the abstract class.

The constant number information class contains two vectors for the double and integer values.  The add element function just extends the two vectors by one element.  The set element function copies the token double and integer values into the respective vectors for the element specified.  No clear element function was needed since there is nothing to clear.  Two array access functions were implemented to access the data in the two vectors, which will be used at run-time.

The constant string information class contains a vector of string instance pointers.  The add element function appends a pointer to a newly created string instance to the vector.  The set element function copies the token string into the element specified.  The clear element function clears the string for the element specified.  Once the string instances are created, they will be reused if dictionary entries are removed.  A destructor was implemented to delete all of the string instances.   An array access function was implemented to access the data in the vector, which will be used at run-time.

Information instance pointers were added to the program model class.  These instances are created in the constructor and passed to their associated information dictionaries.  Both the constant number and string dictionaries are now information dictionaries.  The information instances are deleted in the destructor.  There are no longer any known memory issues.

[commit b9772d4149]

Information Dictionary – New Design

The original design of the information dictionary made the assumption that the additional information would be contained in a vector and was given a structure for the information.  The definition was a class template where the information structure was the argument, which was put into the vector, which the information dictionary handled directly.  However, in the case of the constant number dictionary, two vectors are needed so that memory is not wasted (see last post).

The details of the additional information need to be separated from the information dictionary.  In other words, the information dictionary should not know (or assume) that the additional information is a vector.  An abstract information class can be used that defines the interface to the additional information.  The interface requires several functions for accessing the information:
add element - add a new element to the end of the additional information

set element - set an element from information in a token used when a new element is added or an element previously deleted is reused

clear element - clear the contents of an element when the dictionary entry is removed (made available for reuse)
The actual information classes are derived from the abstract class and implement these functions to manipulate their information as required, which could be stored as a vector, two vectors, or something completely different.  The information dictionary has no knowledge of the information class internals and simply uses the interface functions.

The information dictionary class can be a normal class derived from the dictionary class containing a reference to the additional information.  The abstract information interface functions are used to manipulate the additional information.  The information dictionary needs re-implement these functions from the base dictionary class:
add - adds a new dictionary entry and additional information if not already in the dictionary and returns its index

remove - removes the additional information if the dictionary entry was removed

Friday, November 29, 2013

Information Dictionary Issues

The information dictionary class extended the base dictionary class by adding a vector for additional information and was implemented as a class template (see post from October 6).  The additional information in the constant string dictionary contained a pointer to a string instance (see post from October 6).

The memory leak in the constant string dictionary was caused by how the information dictionary template and constant string information classes were implemented.  The problem occurred when a string in an entry of the information vector was replaced with the same string.  A new information instance was created with a new string pointer, which was put into the information vector, and the old string instance was lost (a memory leak).

When an old program line was dereferenced after the new replacement line was encoded, a string being replaced by the same string had its reference incremented in the dictionary from one to two by the encode, then the dereference decremented the count back to one.  However, when the dereference was moved to after the encode, the reference count of the string went from one to zero and the dictionary entry was freed, but not the entry in the information vector.  When the new line was encoded, a new string instance was created overwriting the old string instance pointer.

While this problem was not difficult to correct, another issue was discovered, this time with the constant number dictionary where its additional information consisted of a double value and an integer value contained in a structure (see post from October 6).  Each double value was aligned on a double boundary (eight bytes) and because an integer is half of a double (four bytes), four bytes of padding is inserted by the compiler between each element in the vector (wasted memory).

The only way to correct this is to separate the two sets of values by having a double value vector and an integer value vector.  Unfortunately, the information dictionary template class only allows for a single information vector.  A new design is needed for information dictionaries.

Program – Dereferencing Replaced Lines

When a line is replaced, references to dictionary entries in the old line must be removed.  This was taking place after the replacement line was encoded.  When a dictionary entry is dereferenced, the reference may no longer be used causing the dictionary entry to be made available for another entry.  The new line may add new dictionary entries, but with the encode before the dereferencing, the new entry will be added to the end of the dictionary if there are no free slots.

It is desirable for new dictionary entries to use slots that may be freed with the old line being replaced.  This will help the dictionary from growing larger then it needs to be.  Therefore, the dereference call was moved to before the encode call.  With this change, the results for encoder test #2 changed slightly, but only with respect to indexes of a couple of dictionary entries.

A previously undiscovered memory error was reported on encoder test #2 when running the memory test script.  The problem occurred in the constant string dictionary with the allocation of the string pointers for the QString instances.  While investigating this problem, another issue was discovered in the constant number dictionary, though this issue is much less serious and only results in wasted memory.  The conclusion was that the information dictionary class (currently defined as a template) needs to be redesigned.

[commit f284a33ac8]

Tuesday, November 26, 2013

Program – Saved RPN Lists (Removal)

The translated RPN lists for program lines were being saved in the line information list, which also contains the offset of the line within the program code, the size of the line and the index to the error list if the line has an error.  These lists were originally used as the source of the data for the program view and were also used to detect when lines changed (by comparing the translated RPN list of the new line with the saved RPN list).

The use of the RPN lists for the program view was removed when the program code array was implemented.  With the decode routine, the program code of the lines can now be converted to an RPN list.  The line change detection in the update line routine was modified to decode the code of the line for the new line being changed to an RPN list, which is compared to the translated RPN list of the new line.  With the RPN lists in the line information list no longer being used, this variable was removed along with all references to it.

A problem was found when RPN lists were compared.  The strings of tokens other than REM commands, REM operators and string constants (whose strings comparison must be case sensitive) should use a case insensitive string comparison.  However, the case sensitivity argument was not supplied to the compare call, so the default comparison used was case sensitive.  The correct argument was added.

[commit dd3e4b62ed]

Sunday, November 24, 2013

Program – Decoder

Before the internal code of a program line can be recreated, the program code needs to be decoded into an RPN list.  Like the encoder, which is part of the program model class because it needs access to the dictionaries, the decoder will also be part of the program model.

The decode routine is given the line information of the line to decode containing the offset of the line within the program code and its size.  A new RPN list is created and for each program word in the line, a new token is created and assigned the code and sub-code of the program word.  If the code has an operand text function in its table entry (implying the code has an operand word), the operand text function is called to get the text for the token from the operand, which is assigned to the string of the token.  The token is added to the RPN list.  After all the words of the line are processed, a pointer to the RPN list is returned.

Like the encode routine, the decode routine is a private function within the program model class.  To access recreated lines of the program, a new line text routine was added.  This routine is given the index to the line and starts be retrieving the information for the line, which is passed to the decode routine.  The pointer to the RPN list returned is passed to the recreate routine of the recreator instance (which was added to the program model class).  The RPN list is deleted and the string returned from the recreate routine is returned.

The temporary check to prevent encoder test files from being used with the recreate output option (-to) was removed from the tester class.  In the tester run routine for encoder test files after outputting the code of the program and the dictionary entries, if the recreate output option was selected, each line of the program is output using the line text routine.

The expected results files for the three encoder tests were created from the encoder test results files with the output of the program added to the end.  All the encoder tests are recreated correctly.  The test script and batch files were updated to also test the encoder test files with the recreate output option.

[commit 1f24a70152]

Saturday, November 23, 2013

Recreator – RPN Lists (Tagged)

The recreator is now complete and all of the initial set of commands (LET, PRINT, INPUT, and REM) are fully supported.  The repository has been tagged v0.6.1 to mark this milestone.  Some minor issues were also corrected for this commit including:
  • Corrected an expected results for the recreated translator test #12 (INPUT tests).
  • Reorganized the access functions in the recreator class.
  • Renamed the recreator class is empty access function to output is empty so as not to be confused with the stack is empty function.
  • Renamed the recreator class last access function to the more explicit output last character and removed the output string is empty check.
  • Corrected some comment formatting and added some missing comments.
  • Corrected formatting issues where spaces were incorrectly added instead of tab characters (problems are caused by QtCreator editor bugs where is doesn't always pay attention to the spacing/tab settings).
  • Added some missing FLAG option comments (details below).
The next major step is to integrate the recreator with the GUI, but in order to do this, the program code of lines needs to be decoded into an RPN list of tokens for the recreator.  When a line is entered into the program, it will be recreated and the recreated text will replaced the entered text in the edit box.  Preferences will be added to control the formatting of the recreated output, for example, if spaces should be added around operators, after commas, etc.  The FLAG option comments show where checks for these options will be.

[commit 0cd7b84700]

Recreator – Colons

Colons between statements are handled by setting the colon sub-code of the last token of a statement.  To support the recreation of colons, a check was added to the main recreator loop after the check for the parentheses sub-code that if the colon sub-code is set, a colon and a space is added to the output string.

This change caused a slight problem with the recreation of the remark operator if there is a colon just before the remark.  Two spaces were being added in front of the "'" operator.  To prevent this, a check was added after a non-empty output string check to also not add a space if the last character in the output string is already a space.

A new last access function was added to the recreator class that returns the last character in the output string or a null character if the output string is empty.  The expected recreated outputs for translator test  #16 (colon tests) were updated and are recreated correctly.

[commit fc90cf0e32]

Recreator – Remarks

There are two codes for remarks, the command code (for the REM command) and the operator code (for the "'" operator).  The token contains the string of the remark.  Generally, the keyword (REM or "'") along with the remark string is added to the output string.  However, there were two issues to be handled.

Unlike all the other commands, no space is required after the REM command when followed by a letter, so a statement like "REMARK A Comment" is valid.  The issue is for a REM statement entered as "remark a comment" in all lower case.  When recreated, the result would have been the "REMark a comment" statement.  So, if the first character of the remark string is lower case, the REM keyword is converted to lower case.  This still won't work if something like "Remark A Comment" is entered.

The second issue involves the remark operator.  A space is needed before the "'" operator if the command is not at the beginning of the line to provide some separation between the previous statement.  To determine if the remark is at the beginning of the line, an is empty access function was added to the recreator class  that returns whether the output string is empty (which implies this the beginning of the line if nothing was added yet).

A single rem recreate function was added to handle both remark codes and a pointer to this function was added to the remark code table entries.  To test the lower case check, an all lower case remark statement was added to translator test #15 (REM tests).  The expected recreated outputs for this test were updated and are recreated correctly.

[commit 03c73c3de7]