18/09/2024
A Syllabic Cipher of Cardinal Gualterio Reconstructed Manually
The letters record real-time reactions to historical events such as James Edward's expedition to Scotland (December 1715), Prince Eugene's victory over the Turks (August 1716), Spanish invasion of Sardinia (August 1717), and the Triple Alliance (January 1717). Actually, these are mentioned in cleartext. Hopefully, the ciphertext contains even more interesting contents. (I don't know whether the plaintext is in the archives.)
02/09/2024
Parliament's Ban on Ciphers during the English Civil War
But the ban was only selectively enforced, at least in the view of the royalists. The Mercurius Aulicus (21 October 1643), a weekly royalist news pamphlet, accused the parliamentarians of the partisan application of the rule and printed an intercepted cipher letter subscribed by a Parliamentarian, a Matthew Durbun, pointing out that the parliamentarians "when they please can practice it, without the least transgression of their order, which it seems was made only for the punishment of the Kings friends but not for such innocent Rebels as they are."
How long the ban was in effect is not known for sure. While it was natural for the royalists to continue using ciphers (see "King Charles I's Ciphers"), we know that under John Thurloe (head of intelligence from July 1653), informants used ciphers (see "Codes and Ciphers of Thurloe's Agents"). Akkerman points out John White's A Rich Cabinet with Variety of Inventions (1653) promoted use of cipher when writing love letters and John Cotgrave's The Wits Interpreter (1655; 1662; 1671) described one of Cardinal Richelieu's cipher keys and recipes for secret ink. By these days, ciphers as well as steganographic techniques such as secret ink became quite common even among ordinary people.
01/09/2024
Duke of Nevers' Variable-Length Figure Cipher (1571)
Although the ciphertext is relatively short, once the continuous stream of figures can be broken into tokens (cipher symbols), homophonic solvers may readily decipher this.
15/08/2024
A Second Copy of Music Cipher to Charles II Acquired by the British Library (ca.2018)
What I didn't know then is that the cipher to Charles II quoted in my article is in the British Library (Add MS 45850, f.68), but its provenance through the Port family I found in googling is about another copy, which the British Library newly acquired (as of February 2018) (Add MS 89288). I learned of this in the British Library's blog article, "'Conceal yourself, your foes look for you': revealing a secret message in a piece of music" (20 February 2018).
I was reminded of this topic when reading Nadine Akkerman and Pete Langman (2024), Spycraft, which refers to Nadine Akkerman (2018), Invisible Agents (the BL's blog also mentions forthcoming publication of this book). Akkerman discusses the nineteenth century copy in BL Add MS 45850 and considers it a hoax. Her dismissal of Jane Lane as the author based on the latter's literacy level is convincing. Even if we assume other authorship, it is hard to think of circumstances in which this kind of cipher came into play ("it is not as if Charles Stuart did not know his enemies were searching for him" etc.).
14/08/2024
Use of Diacritics/Exponents/Vowel Indicators in Milanese Ciphers
I now see Milanese ciphers used combined symbols with diacritics or exponents as early as the mid fifteenth century but syllables were not formed systematically as with vowel indicators. I now added a section about this: "Use of Exponents/Diacritics/Vowel Indicators in Milanese Ciphers (1450s-1530s)."
It seems systematic vowel indicators are a degenerated form of such ciphers, but dating of the ciphers in the archives is necessary to assess such a hypothesis.
13/08/2024
Constantijn Huygens Jr.'s Secret in Simplistic Concealment Cipher
Constantijn Huygens Jr. (1628-1697), a brother of the physicist Christiaan Huygens, used a cipher in some part of his journals. In late twentieth century, it was found out that encrypted words can be read by simply ignoring odd-numbered letters. For example, b.mregtvelennphnōdesr reads met een hoer ("with a prostitute") (there is one extra letter, which may be an error). The cipher typically concealed such embarrassing privacy of the diarist. I learned of this in Christopher Joby (2014), The Multilingualism of Constantijn Huygens (1596-1687) p.282.
Constantijn Jr. was secretary of Prince of Orange William III (my favoutie historical character) since the latter became stadtholder in 1672. He records his personal experience in participating in major events such as William's campaings to oppose the French invasion, the expedition to England (the Glorious Revolution), and the Irish campaign to prevent the return of James II. Joby (2014) includes some quotes from these (p.284 ff.), and more would be found in Rudolf M. Dekker (2013), Family, Culture and Society in the Diary of Constantijn Huygens Jr, Secretary to Stadholder-King William of Orange. His journals seem interesting in describing historical events from a personal point of view. For example, the day after William III's coronation, from which he was absent, he wrote, "In the early afternoon I was with the king, who asked me where I had watched the coronation. I said that I had been busy deciphering the resolution of the States General, received in cipher, about the alliance with the Emperor, because I thought that the king would want to read this quickly. He asked me if I had received a coronation badge, and I answered no, without receiving much of a response." (ibid. p.40)
12/08/2024
Codebreaker Constantijn Huygens
I've heard that Constantijn Huygens Sr. (1596-1687; the father of the famous physicist, Christiaan Huygens) did codebreaking, but it was only recently that I learned that he regularly served in that capacity in Chapter 2 of Nadine Akkerman and Pete Langman (2024), Spycraft (p.153-156). According to this book, he studied cryptanalysis at the University of Leiden in 1616 and even got a pay raise while serving as a secretary to Prince of Orange Frederick Henry since 1624. The authors translate his proud words about his achievements in his autobiography: "At every siege, I proved my skills, anticipating the tricks of the enemy by means of my own knowledge of deceit ...." Particular reference was made to his contribution to the siege of Breda (1637) when requesting a pay raise.
His first achievement in the field appears to have been during Frederick Henry's siege of 's-Hertogenbosch in 1629, when he was asked to decipher intercepted Spanish letters in cipher by using his knowledge of Spanish (Christopher Joby (2014), The Multilingualism of Constantijn Huygens (1596-1687), p.78).
He was not always successful. At one time in 1634, he said ciphers of the king of Spain were "more difficult to conquer than" the king himself. (Akkerman and Langman, p.154)
While his library contained many books on cryptography, ciphers he designed for royalists during the English Civil War were simple homophonic substitution ciphers albeit with an extensive nomenclator (ibid. p.156). (I'm inclined to think such ciphers were the most practical after all. John Wallis also proposed a simple Caesar cipher when asked for an "easy cipher", as noted in "John Wallis and Cryptanalysis".)
(By the way, Joby (2014) discusses "code switching", which has nothing to do with cryptography and may be broadly understood as switching to different languages when quoting etc.)
11/08/2024
A Second "More Ample" Babington Cipher
Chapter 2 of Nadine Akkerman and Pete Langman (2024), Spycraft details on the undoing of Mary, Queen of Scots, in the Babington Plot and points out many facts which should have been apparent from well-known sources.
(1) The famous cipher used in Mary's fatal letter of 17 July 1586 to Anthony Babington was a very simple one. (The first letter from Mary to Babington dated 25 June 1586 (Pollen (1922) p.15) was also in the same cipher, as I noted in "Ciphers of Mary, Queen of Scots". The 17 July letter was a reply to Babington's response to this first letter.) When one learns that Mary used more elaborate ciphers with other correspondents (as seen in the collection of keys in SP53/22, SP53/23), one cannot help wondering why such a simple cipher was used in this important correspondence.
The authors point out (p.133) that new correspondents were given a simple cipher at first and a fuller one later. Indeed, the 17 July letter ends with "I have commanded a more ample alphabet to be made for you, which herewith you will receive" (cf. Pollen (1922), p.26 ff., esp. p.45).
According to the authors (p.138), such a "mature" cipher was actually used by Babington in his reply dated 3 August 1586 (Pollen (1922) p.46-47, printed from SP53/19 no.10), but the key was intercepted by the authorities and the letter was readily deciphered by Phelippes. Phelippes attests "The new Alphabet sent to be used in time to come between that Queen and Babington ... is of Nau's hande" (I have not checked the cited SP53/19 no.85, which is Phelippes' record of the secretaries' testimony from 4 September, to see whether "in time to come" really refers to the 3 August letter).
[(8/14/2024) Nau's argument is printed in Tytler's History of Scotland.
It says "The new alphabet sent to be used in time to come between that Queen and Babington, accompnying the bloody despatch, is of Nau's hand." So the cipher was attached to the 17 July letter (consistent with Mary's wording "herewith") and it may well have been used in the 3 August letter. Curle's cover letter to Barnaby enclosing the 17 July letter says "Giuen hereiwth is the addition to this alphabet" (Pollen (1922) p.25). If this does not refer to an update to the Mary-Babington cipher, such an update may well have been enclosed at the same time.]
[(5/31/2025) According to Alford, Watchers, "A proposed new cipher between Mary and Babington is set out in BL Additional MS 48027 f.313v." (Note to Chapter 14)]
(2) Another observation of the authors interesting from a cryptologic viewpoint (p.134-135) is that the well-known copy of the Mary-Babington Cipher in SP12/193/54 is a copy from the original, rather than a product of Phelippes' codebreaking. This can be seen from the nomenclature entries such as "your name" and "myne."
On the other hand, before concurring with the authors' conclusion that Phelippes did not break the Mary-Babington cipher at this time because he already had the key, we may need to assess whether it was feasible to associate a particular intercepted key to Babington (for example, if the three keys on SP12/193/54 were the only keys sent out around June 1586, the authentic key could have been used by Phelippes).
10/08/2024
Codebreaker John Somer in the 1580s
(2024/09/08) I remembered my article on Mary's ciphers mentions Walsingham's forwarding Mary's letter to be deciphered by Somer in October 1582. I now added a mention of this.
22/07/2024
Bolton's Telegraph Code (1871) Adopted by Japanese Foreign Ministry
I added this note to "Nonsecret Code: An Overview of Early Telegraph Codes" and "日本の電信暗号".
According to John McVey's website, a copy of the codebook is in the British Library.
21/07/2024
Three Unknown Chinese Codebooks (ca. 1905)
19/07/2024
Ciphertext-only Attack on Classical Ciphers by Using AI
My interest is in a ciphertext-only attack on homophonic ciphers or syllabic ciphers, which are the most common ciphers in the early modern period. Although I have not found a work on neural network approach on these, two (also cited by Closa (2023) mentioned in the previous post) are interesting in dealing with ciphertext-only attack on the Caesar (shift) cipher, the Vigenere cipher, and (monoalphabetic) substitution cipher.
Focardi et al. (2018) claims to be the first to provide a ciphertext-only attack on substitution ciphers based on neural networks (Abstract). It assumes a weakness of the cipher is given and the neural network exploits it.
(i) In the case of the Caesar cipher (a.k.a. the shift cipher), the key is a single number (e.g., up to 26). The frequencies of symbols generally reflects the frequency of letters (e.g. of English-language text). Thus, a neural network is trained to predict a key from the frequencies. The trained neural network can "recognize" the key without actually trying shifts to see whether readable text is obtained.
(ii) For the Vigenere cipher, a key length (m) is assumed and the Caesar classifier is applied to subtexts composed of symbols taken from the ciphertext at distance m. This is like a conventional method by using brute force but employs the Caesar classifier from (i).
(iii) For general (monoalphabetic) substitution ciphers, an overall framework is similar to conventional hill climbing. Starting from a random key, a "goodness" value is computed. After changing the key a little bit, the "goodness" value is computed again. If it has improved, the change is kept; otherwise the change is discarded. By iteration, the key is improved step by step. Again, the authors' method does not need to actually try intermediate keys to see whether a plausible plaintext is obtained. Instead, a neural network is trained with 3-grams of both plaintexts and ciphertexts. Although they "have no guratantees that this will provide a good classifier which is able to tell 'how' similar is a text to a plaintext and, consequently, how good is a key, but experimental results have confirmed that the method is effective." It will be interesting to see whether this works for homophonic ciphers or syllable ciphers.
Ahmadzadeh et al. (2021) seems to achieve what I thought impossible: training with plaintext/ciphertext pairs (Table 4) allows deciphering a ciphertext with an unknown key. The decryption function is learned "regardless of the cipher complexity or key length" (IV D). While experiments were done with Caesar, Vigenere, and (monoalphabetic) substitution, the authors consider their approach "has the potential to crack modern ciphers" (IV H).
To learn the decryption function from the plaintext/ciphertext pairs, they used an attention-based LSTM encoder-decoder model (Fig. 3). To quickly recall the terminology, a class of neural networks that can have "memory" is a recurrent neural network (RNN). A class of RNN that solves its problem (vanishing gradients in deep networks: "A small gradient value does not contribute very much to learning") is LSTM. A problem with LSTM ("a lengthy input sequence causes LSTM to forget important information along the sequence") is solved by an attention mechanism, which allows "dynamically highlighting important features of the input data" (III A).
As a prerequisite, they assume the ciphertext has "punctuation" and can be readily parsed into words (Fig. 4). It will be interesting to see whether their work can be generalized to ciphertext without punctuation.
References:
Riccardo Focardi and Flaminia L. Luccio (2018), "Neural Cryptanalysis of Classical Ciphers", Italian Conference on Theoretical Computer Science (ICTCS 2018), Urbino, Italy, September 18-20, 2018 (CEUR-WS.org).
Ezat Ahmadzadeh, Hyunil Kim, Ongee Jeong, and Inkyu Moon (2021), "A Novel Dynamic Attack on Classical Ciphers Using an Attention-Based LSTM Encoder-Decoder Model", IEEE Access, 2021, vol.9, pp.60960-60970, DOI: 10.1109,ACCESS.2021.3074268 (IEEE Xplore)
13/07/2024
Codebreaking with AI
For a long time, I couldn't find a positive answer to this question. Although my search was half-hearted, even Copilot with Chat-GPT-4 did not give a relevant answer to my question: "Are there papers about codebreaking by using AI?" Then, I noticed a poster abstract in HistoCrypt 2024: Oriol Closa, "Polyalphabetic cipher decryption function learning with LSTM networks", which seems to be based on her master thesis, Closa Oriol (2023), "LSTM-attack on polyalphabetic cyphers with known plaintext: Case study on the Hagelin C-38 and Siemens and Halske T52" (KTH). It teaches that "the application of Machine Learning to extract key information from intercepts is not a well researched area yet." (Abstract) and there are even "many authoritative opinions within the field" against utility of machine learning in classical cryptography (p.57).
Machine "learns" by finding a best parameter set for a model, which is like a very complex filter that receives an input and produces an output. In an example of machine translation, the input is a sequence of words in English and the output is a sequence of words in Japanese. In order to train a machine in this example, bilingual corpus of corresponding texts in the two languages is fed to the computer, whereby the computer learns an English-Japanese translation model. Given a new text in English, the trained computer can apply the model (filter) on the input to produce a text in Japanese as an output.
By analogy, given a corpus of ciphertext/plaintext pairs, AI may learn to decipher a new ciphertext into a plaintext. However, it should be limited to the case where the new ciphertext is based on the same cipher key used for training -- that's what I thought. But the thesis taught me machine learning can do more than that. (The key is included in the training data (p.39).)
Remember that the output of a trained model need not be similar to the input. For example, when the input is a text in English, the output may be some classification or labelling of the text rather than a text in another language. In Oriol (2023), in my understanding, the input is ciphertext, a crib (known plaintext, presumably corresponding to the ciphertext), and a null key (a placeholder in the input vector), and the output is plaintext (which should match the crib and is included only for analysis) and the extracted key (p.34). Thus, this seems to receive a maching ciphertext/plaintext pair to produce its key ("extract the external key given a combination of plaintext and ciphertext without the use of the internal setting" p.51).
The thesis deals with four ciphers: Vigenere, Playfair, Hagelin C-38, and Siemens and Halske T52 with LSTM networks (a kind of neural network). Main differences among them are the decipher function reflecting the cipher scheme (I guess this means the cipher algorithm is known and only the key is to be found out), the crib length (e.g., 15, 25), and the size of the hidden layer (e.g., 256, 2048) (p.33, 40). The author says LSTM networks can extract key information given a crib (in my understanding, this is a matching plaintext/ciphertext pair) for Vigenere, Hagelin C-38, and Siemens and Halske T52, but not Playfair.
(16 July 2024) Ajeet Singh, Kaushik Bhargav Sivangi, and Appala Naidu Tentu (2024), "Machine Learning and Cryptanalysis: An In-Depth Exploration of Current Practices and Future Potential", Journal of Computing Theories and Applications (JCTA, DOI: https://doi.org/10.62411/jcta.9851), Vol. 1 No. 3 (2024) also says "the integration of machine learning, and specifically deep learning, into cryptanalysis has been relatively unexplored."
12/07/2024
George Lasry's Paper on Syllabic Cipher Released
Details of his algorithm are now published in George Lasry, "Deciphering Historical Syllabic Ciphers" (HistoCrypt 2024).
One might be tempted to say extension of a homophonic solver to a syllabic solver only involves enlarging the search space for, say, 25 letters of the alphabet by adding variables for 25*25=625 syllables. However, since last year, I've been wondering how to design a scoring function for intermediate decipherments. In typical homophonic solvers, 4-grams or 5-grams are used in computing a score. For example, if an intermediate decipherment abounds in plausible 5-grams (e.g., "ISION", "EMENT", "ETLES" for the French language), it receives a high score. But when symbols represent syllables such as "SI", "ON", "EM" "EN", just random assignment of these syllables might result in a relatively high score.
The answer to my year-long question turned out to be very simple. The scoring function is based on 4-grams, composed of not only single letters but also syllables. Thus, a 4-gram may be "E-S-T-A" or "CO-N-TRO-L" (p.177).
While it is easy to say this, it is quite another to implement it, because this involves re-constructing the whole language model. A language model represents frequency characteristics of all possible 4-grams (or 5-grams etc.) in a large corpus of text in the language in question. For a language model for conventional homophonic solvers, creating a language model amounted to simply counting letter groups. For a syllabic solver, it is necessary to first break text in the corpus into syllables, but there can be many ways to split a word. For example, the word ESTABLISHED may be broken into E-S-TA-B-LI-S-HE-D (using only CV syllables), ES-TA-B-LI-S-H-ED (using CV and VC), E-STA-BLI-S-HE-D (allowing also CCV), or the like (p.174). Naturally, success of the algorithm depends on the choice of the set of syllables used and the decomposition scheme, which required "extensive trial-and-error to fine tune" (p.176, 179).
And of course, the more than ten-fold increase in the number of variables (which more than exponentially expands the search space) means "extensive computing power" is required. George used a 64-core Windows 10 Pro PC with 256 Gbytes of RAM memory (p.176), whereas one commercial PC I see on the web has 4-core and up to 3.40 GB of memory.
His groundbreaking new scheme has proven its merits by breaking ten ciphertexts (including five with known keys).
04/07/2024
Reconstructed Ciphers related to Mary, Queen of Scots, preserved in Scottish Catholic Archives (SCA)
03/07/2024
Early French Figure Ciphers
It is used in a letter from Henry IV to Duke of Nevers, 12 September 1592 (BnF fr.3620, f.70-71)(p.47). The cipher employs variable-length figures written continuously, but the authors could parse the ciphertext into cipher symbols by assuming three-digit figures always start with "1" (p.48; see also my articles from 2017-2019 and 2018-2019). After the key was reconstructed, the authors found the original cipher table among Nevers' collection of keys in BnF fr.3995, fol.140 (p.49).
Many cipher letters between Henry IV and the Duke of Nevers have been known but they used conventional symbol ciphers in BnF fr.3995, fol.67 rather than numerical ciphers (p.49-50, 46; see also my article). The authors point out that the 1592 letter in question is countersigned by Martin Ruzé de Beaulieu, while all the other letters in cipher from the king to the duke from 1591 to 1594 are counstersigned by Louis Potier de Gesvres (p.49-50). The 1592-cipher was also used in letters in 1591 to Henry IV (two from Duke of Biron, one from sieur de Guitry, baron de Salagnac, and marquis de Pisani) (p.51).
The paper is also valuable in citing many examples of exclusively digit ciphers in the 1580s/1590s (p.50; see also other figure ciphers mentioned in my articles: Henry III etc., Buzenval).
(5 July 2024) I now remembered another French ciphertext in variable-length figures (1, 2, or 3 digits) written continuously (ca.1620?) is discussed in here.
(24-25 August 2024) The paper points out at least 23 letters sent to the Duke of Nevers in 1589-1591 use only digits (p.50). To this list may be added another ciphertext in digits continuously written without break (since it is not deciphered, we cannot tell whether it employs variable-length symbols). The ciphertext is in an Italian letter from Lodovico Birago to the Duke of Nevers, 13 November 1571 (BnF fr.3251, f.119), which I mentioned here some years ago.
01/07/2024
Cross-Cipher Errors - A New Modality of Communication Analysis
23/06/2024
Some Updates on Correspondence between Philip II and Vargas Mexia
I also found further undeciphered letters from Vargas Mexia are catalogued in BL Add MS 28421.
I added these under the section "(Additional Note, 23 June 2024)."
09/05/2024
Early Japanese Syllabary Table in Milanese Archives
25/04/2024
Korean Telegraphic Code
Once characters are encoded into digits or roman letters, encryption methods including substitution and transposition are applicable. Today's computer can of course handle Hangul characters. But in the early years of telegraphy, telegrams in Korea had to be in Chinese characters or Latin letters.
So, a telegraph codebook for Chinese characters were used in Korea.
I have seen a Korean version (『漢電』) of a Chinese telegraph codebook.
Even after WWII, it appears a telegraph codebook with similar content was used, in view of an edition adapted for use by those who could not read Chinese characters (Korean Telegraphic Code Book, with characters arranged by sounds in English alphabetic order according to the McCune-Reischauer system of transliteration).
These are already covered in 電碼――中国の文字コード, dealing with Chinese telegraph codes, in which I now made small corrections.
(By the way, I wrote the above ten days ago, but I couldn't upload it because my smartphone failed and I couldn't pass the two-factor authentication for logging into the blog.)